{"id":"56323680-b83a-44cb-842b-11604e13c75b","arxiv_id":"2507.19082","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a VR kitchen, people delegated tasks and placed robots differently based on robot morphology, yielding three hypotheses about biomorphism, perceived sensing, and robot bulk.","lead":"This paper reports a virtual reality kitchen study in which 22 participants arranged cooking items and a robot for a shared meal, and it uses their placements and spoken thoughts to form three hypotheses about robot design. It matters because it gives robot designers concrete, testable guesses about how body shape changes trust and task delegation in collaborative cooking.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Figure 5's rank ordering rests on ~4 participants per robot, so the grounding of H1–H3 may be weaker than claimed.","rationale":"The reader's weakest assumption concerns external validity: whether VR arrangements with static robots generalize to real kitchen collaboration. That is a real limitation but one the paper explicitly acknowledges. My concern is more load-bearing because it targets the internal grounding of the hypotheses themselves: the quantitative patterns in Figure 5, which motivate H1 and H3, are based on very small per-robot samples with no inferential statistics. If those patterns are not robust, the central contribution—'hypotheses grounded in observed participant behavior'—is weakened even within the study's own scope. The paper is honest about being exploratory, and the reader's CONDITIONAL verdict already flags the small sample and missing statistics, so I do not propose changing the verdict. However, the specific test I propose would settle whether the observed rank ordering is stable or an artifact of sampling and confounds. This is a partial agreement with the reader: we both see the sample-size issue, but the reader's stated weakest assumption focuses on VR-to-real transfer rather than the internal evidentiary basis of the hypotheses.","tokens_in":8376,"tokens_out":4395,"duration_ms":50101,"concrete_test":"Re-analyze the task-assignment data from Section IV-A using bootstrap resampling: for each robot, resample the 3–5 participant outcomes 10,000 times and compute the 95% CI for the collaborative-task mean and the collaborator ratio. Then compare the biomorphic versus technomorphic robot groups with a permutation test. If the 95% CI for the group difference includes zero, or if the rank ordering of Pepper/Spot above Stretch/PR2 appears in fewer than 50% of bootstrap samples, the claimed grounding of H1 and H3 is not statistically supported. Additionally, repeat the analysis excluding Spot to test whether the biomorphic preference is driven by prior familiarity with this commercial robot.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the analysis grounds three hypotheses (H1–H3). The quantitative evidence in Section IV-A (Figure 5) is the main support for H1 and H3, but each robot was encountered by only 3–5 participants (22 participants, 11 robots, 2 recipes each). The paper reports no confidence intervals, no inferential tests, and no control for robot-specific confounds such as prior familiarity (Spot is a widely known commercial robot), color, or brand. The qualitative annotations that support H2 were coded primarily by a single annotator after intercoder reliability was checked on only 7 participants, with κ=0.62 for the agent-placement category (Section III-D). Given these small, unbalanced samples, the rank ordering in Figure 5 is fragile: a single participant's 'collaborate or not' decision shifts the collaborator ratio by 0.25. The paper acknowledges the data are 'too small to extract statistically robust conclusions' (Section IV-A), yet the hypotheses are presented as grounded in observed participant behavior. The concern is that the apparent preferences for Pepper/Spot and the avoidance of bulky robots may be artifacts of which specific robots were sampled and which participants saw them, rather than robust effects of morphology.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an exploratory VR study in which 22 participants arranged a virtual kitchen for collaborative cooking with one of 11 robots (two recipes per participant, random robot assignment), producing 3D placements, think-aloud protocols, and interview transcripts. The authors analyze task assignment frequencies per robot (Figure 5) and qualitative annotations of the think-alouds. Based on these observations, they formulate three hypotheses: H1 (more biomorphic morphology encourages task sharing), H2 (beliefs about sensory capabilities are inferred from action capabilities suggested by the body, e.g., 'if it has a hand it can probably feel by touch'), and H3 (gracile robots are avoided less than imposing ones). The paper explicitly frames the contribution as hypothesis generation, not hypothesis testing, and states that follow-up studies are needed to verify the hypotheses.","tokens_in":8578,"tokens_out":3038,"duration_ms":32912,"significance":"If taken as an exploratory hypothesis-generating contribution, the paper has real value: the mise en place scenario in VR is a relatively novel and ecologically plausible method for studying perceived affordances from robot morphology, the dataset is openly available, and the hypotheses H1-H3 are concrete, falsifiable, and clearly separated from confirmatory findings. The authors are commendably transparent about the exploratory nature of the work and about limitations such as static robot models and small sample size. The use of a structured annotation scheme with reported inter-coder reliability (albeit limited) and the explicit linkage to the MetaMorph taxonomy are strengths. However, because the central claim is that the hypotheses are 'grounded in observed participant behavior,' the quantitative and qualitative evidence base must be robust enough to support that grounding; the current evidence is fragile in several load-bearing places.","major_comments":[{"comment":"The rank ordering in Figure 5 that motivates H1 and H3 is based on only 3-5 participants per robot (22 participants, 11 robots, 2 recipes each, no pair repeated). The paper itself says the data are 'too small to extract statistically robust conclusions' (Section IV-A), yet the hypotheses are presented as grounded in the observed patterns. A single participant's 'collaborate or not' decision changes the collaborator ratio in Figure 5c by about 0.25 for a robot seen by 4 participants. No confidence intervals, bootstrapped errors, or inferential tests are reported, and no account is taken of robot-specific confounds such as prior familiarity (e.g., Spot is a widely known commercial robot), color, or brand. To make the claimed grounding credible, the authors should either report per-robot n, raw counts, and some measure of uncertainty (e.g., bootstrap CIs), or explicitly re-label the patterns as anecdotal rather than hypothesis-grounding.","section":"§IV-A, Figure 5"},{"comment":"The hypotheses are post-hoc interpretations of the same dataset from which they are derived. The authors state that 'we had to look at other features of the popular robots to identify what they may have in common' (Section V), which is a direct admission that H1 and H3 were selected after seeing the data. This is acceptable for an explicitly exploratory paper, but the current wording ('formulated our hypotheses') could be read as implying more independence than exists. The manuscript should clearly state, for each hypothesis, that it was generated from the observed patterns rather than predicted a priori, and should describe a specific preregistered confirmatory design that would test H1-H3 with a new sample. Without that, the hypotheses are restatements of the ranking in Figure 5 rather than generalizable claims.","section":"§V"},{"comment":"The central support for H3 is the claim about avoidance strategies (e.g., 'participants rarely envisioned sharing proximity with the robot and often took measures to avoid it'), but the annotation data behind this claim is explicitly not shown: 'additional annotations of the TAPs (not shown here due to space constraints).' The numbers given (7 of 22 participants for the first recipe, 3 of 22 for the second) are not linked to robot identity, so the reader cannot verify the proposed association between gracile morphology and fewer avoidance strategies. Since H3 is one of the three central hypotheses, the supporting per-robot avoidance data should be presented (e.g., as a table or additional panel in Figure 5) or the hypothesis should be explicitly downgraded to an impression from the transcripts.","section":"§IV-B"},{"comment":"The inter-coder reliability check is based on only 7 participants (participants 1-8 excluding 4), and the Cohen's kappa for the category that directly supports H2 and H3 (agent placement) is 0.62, which is generally considered moderate agreement. After this check, a single annotator coded all remaining participants. Given that the qualitative annotations carry much of the interpretive weight for H2 and H3 (e.g., judgments about sensoric vs. motoric capabilities and avoidance strategies), the reliability evidence is thin. The authors should either (a) report kappa by subcategory (e.g., safety, obstruction, sensory vs. action inferences) to show which specific codes were reliable, or (b) present a second-annotator pass on a random subset of the remaining participants. As it stands, the qualitative foundations of H2 and H3 are not sufficiently demonstrated.","section":"§III-D"}],"minor_comments":[{"comment":"The abstract statements 'humans prefer to collaborate with biomorphic robots...' are written as assertions in the first reading; while the following sentence clarifies these are hypotheses, the phrasing should be adjusted to make the hypothetical status explicit (e.g., 'we hypothesize that humans prefer...').","section":"Abstract"},{"comment":"The word 'ifluenced' is a typo for 'influenced' in the sentence 'asked if and how the robot's appearance and size ifluenced its placement.'","section":"§III-C"},{"comment":"Figure 5 is difficult to parse because panels (a), (b), and (c) share the same x-axis ordering (sorted by average task assignments in panel a), but the reader must infer this from the caption text. Please state explicitly in the caption that all three panels use the same x-axis ordering, or sort each panel independently for readability.","section":"Figure 5"},{"comment":"The participant section reports 23 recruited, data from 22 after technical loss; it would be helpful to state the age range and gender split for the 22 participants whose data were analyzed (the current numbers appear to mix the full 23).","section":"§III-B"},{"comment":"The observation that participants gave tasks to robots 'with no clear indication of visual sensors' (e.g., Spot) is interesting, but the term 'no clearly visible cameras' is potentially contested; consider adding a supplementary image or a MetaMorph-based sensor-feature table to make this claim verifiable.","section":"§IV-B"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper for what it is: a clearly labeled exploratory study that produces three concrete hypotheses about how robot morphology affects task delegation and spatial avoidance in a kitchen scenario. The mise en place setup with eleven standardized robots and the distinction between perceived competence and willingness to collaborate are genuinely useful angles. The authors also ship an open multimodal dataset and connect their stimuli to the MetaMorph taxonomy, which makes the work reusable.\n\nThe strongest part is the qualitative analysis. The observation that participants inferred hidden sensory abilities more readily than motor abilities is interesting and grounded in the think-aloud data. The finding that gracile robots were avoided less, even when large, gives designers a concrete handle. The paper earns credit for explicitly saying these are hypotheses, not findings, and for being upfront about the VR limitations.\n\nThe soft spots are real but proportionate. With 22 participants and 11 robots, each robot was seen by roughly four people, so the rank ordering in Figure 5 is fragile. The paper says this itself, which lowers the sting, but it means H1 and H3 rest on patterns that could easily shift with a few more participants. The stress-test note about the omitted annotations is also fair: the numbers of participants who shared proximity (7 and 3) are mentioned in prose but not shown in the figures, so you have to take that on faith. The intercoder reliability was checked on only seven participants and the rest was coded by one annotator; again, acceptable for an exploratory study, but not a strong epistemological foundation. The hypotheses are post-hoc interpretations of the same small dataset, not independent predictions. That is a circularity concern, but the authors frame them as hypotheses for follow-up, which is the right move.\n\nWho is this for? Researchers working on robot appearance perception, especially in domestic or kitchen HRI, and anyone building morphology taxonomies. A design-oriented reader will get heuristics worth testing; a strict methodologist will find the evidence too thin to use directly.\n\nMy verdict: this deserves a serious referee. It is not a definitive result, but it is an honest, well-scoped study with transparent limitations and a useful open dataset. I would send it to peer review for an HRI venue, with the expectation that the authors add the missing annotation details and perhaps tone down the presentation of the quantitative ranking. I would not cite it as evidence for the hypotheses, but I would cite it as an example of an exploratory VR paradigm and for the dataset.\n\nRecommendation: engage with it, but treat the hypotheses as promising leads, not established effects.","headline":"An honest, exploratory HRI study that generates three testable hypotheses from a small but openly shared VR dataset; the quantitative grounding is thin, but the paper never overclaims and the qualitative observations carry the weight.","tokens_in":9085,"tokens_out":1355,"would_cite":true,"duration_ms":17861,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Robot body shape changes which kitchen tasks people trust them with, a virtual-reality study suggests.","keywords":["human-robot interaction","robot morphology","virtual reality","perceived affordances","task delegation","kitchen collaboration","think-aloud protocol","shared spaces"],"falsifier":"A larger follow-up in a physical kitchen with a moving robot, measuring the distances participants keep and which subtasks they delegate to robots of different body shapes, would settle the hypotheses: if bulky or non-biomorphic robots are not given fewer shared tasks, and if slender robots do not draw people closer, the claims would be contradicted.","tokens_in":8211,"feed_emoji":"🤖","tokens_out":8638,"duration_ms":87544,"temperature":0.7,"pith_summary":"This exploratory study asks whether a robot's visible body shape changes how people would work beside it in a kitchen. Twenty-two participants used a virtual-reality kitchen to arrange ingredients, tools, and one of eleven differently shaped robots for a mise en place task, narrating their reasoning as they went. From the resulting placements, think-aloud transcripts, and interviews, the authors put forward three testable hypotheses: people prefer to collaborate with biomorphic robots; people infer a robot's sensory abilities less from its shape than its action abilities; and people use fewer avoidance strategies around slender, less bulky robots. If these hypotheses hold in larger studies, robot appearance could be used to encourage task-sharing and comfortable co-location without changing a robot's underlying software.","feed_headline":"Robot shape changes which kitchen tasks people trust them with","feed_subtitle":"A VR mise en place test with 22 people links robot silhouette to task delegation and proximity tolerance.","key_machinery":"The central instrument is the mise en place scenario: a pre-cooking kitchen-arrangement task conducted in virtual reality, in which participants place tools, ingredients, and a static robot model that they are told to imagine as fully capable. Each robot's morphology was categorized along axes such as silhouette (humanoid-hybrid, zoomorphic, technomorphic), size, and bulkiness, and the arrangements, think-aloud protocols, and post-task interviews were annotated for factors like safety concerns, willingness to share proximity, and who was assigned each recipe subtask. This setup lets the paper measure perceived affordances through spatial decisions rather than through rating scales alone.","core_discovery":"The paper claims that robot morphology acts as a surface for perceived affordances: humans infer from a robot's body what it can do and how it will behave, and they use those inferences when deciding whether to delegate tasks and whether to allow the robot into their personal workspace. Specifically, the authors hypothesize that biomorphic silhouettes invite task-sharing, that action-relevant body parts (grippers, arms, height) shape beliefs about motor competence more than visible sensors shape beliefs about perception, and that 'gracile' robots (slender rather than necessarily small) are avoided less than physically imposing ones. These hypotheses are grounded in observed patterns, such as robots with humanoid-hybrid or animal-like silhouettes receiving more offers to collaborate while being given fewer solo tasks, and participants arranging separate workstations to keep a buffer from bulkier robots.","pith_inferences":["The paper's account of H2 implies a possible over-trust effect: people may assume sensing abilities a robot does not actually have, so designers may need to make sensing limits visible rather than rely on morphology.","The avoidance differences attributed to bulky versus gracile bodies could really be driven by perceived predictability or maneuverability; a follow-up that controls for silhouette while varying only bulk, or that animates a bulky robot with graceful motion, could separate these.","The mise en place protocol could be recycled to test whether the type of tool (sharp versus fragile) interacts with morphology, since the paper's related work suggests knives and fragile objects produce different danger perceptions when handled by a robot.","The study's small sample and static robots leave open whether the biomorphic preference is specific to kitchen collaboration or generalizes to other shared manual tasks, such as assembly or cleaning."],"forward_implications":["If H1 is correct, giving a robot a biomorphic silhouette is a design lever for increasing a person's willingness to collaborate on a task.","If H2 is correct, visible camera and sensor placement may matter less for trust than the shape of arms, grippers, and body size, because people assume hidden sensing.","If H3 is correct, making a robot less bulky, even without making it smaller, should reduce avoidance behaviors and make shared workspaces more acceptable.","Because perceived competence and willingness to collaborate appeared to split apart, evaluating a robot's capability alone is not enough; studies should measure spatial and task-sharing choices separately.","The same protocol can be scaled to a larger participant pool and to moving robots to test whether the hypotheses survive outside imagination."],"supporting_citations":[{"why":"Provides the morphological feature taxonomy used to describe and compare the eleven robots' visual characteristics.","marker":"[6]"},{"why":"Shows that anthropomorphic appearance increases trust in human-robot interaction, motivating the study's focus on appearance.","marker":"[3]"},{"why":"Demonstrates that the shape of individual robot components influences perceived personality, supporting feature-level analysis of morphology.","marker":"[8]"},{"why":"Offers evidence that kitchen objects are perceived as more dangerous when handled by a robot, framing safety-related placement decisions.","marker":"[18]"},{"why":"Supplies the two recipes whose tasks were used in the mise en place scenario.","marker":"[22]"},{"why":"Supports the use of virtual reality as a setting where behavior resembles real-world interactions, justifying the VR methodology.","marker":"[9]"},{"why":"Measures participants' general attitudes toward robots, used as a baseline in the participant sample.","marker":"[23]"}],"fun_headline_variants":["Robot shape steers which kitchen tasks earn your trust","In VR kitchens, robot bodies dictate task delegation","Morphology clues: how robot form wins collaboration","Study: robot silhouette shapes kitchen task divisions","Gracile robots get closer in VR cooking study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that how people arrange a kitchen and delegate tasks in virtual reality, with a robot they are told to imagine as competent, reflects how they would actually behave in a real kitchen with a moving robot.","fun_headline_variants_meta":{"raw":{"variants":["Robot shape steers which kitchen tasks earn your trust","In VR kitchens, robot bodies dictate task delegation","Morphology clues: how robot form wins collaboration","Study: robot silhouette shapes kitchen task divisions","Gracile robots get closer in VR cooking study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00016,"raw_usage":{"total_tokens":1191,"prompt_tokens":864,"completion_tokens":327,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":254}},"tokens_in":480,"tokens_out":327,"duration_ms":4120,"temperature":1.0,"reasoning_tokens":254,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:27:00.810212+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A larger follow-up in a physical kitchen with a moving robot, measuring the distances participants keep and which subtasks they delegate to robots of different body shapes, would settle the hypotheses: if bulky or non-biomorphic robots are not given fewer shared tasks, and if slender robots do not draw people closer, the claims would be contradicted.","supporting_citations":[{"cited_title":"Metamorph – a metamodelling approach for robot morphology,","cited_arxiv_id":null,"evidence_quote":"Provides the morphological feature taxonomy used to describe and compare the eleven robots' visual characteristics."},{"cited_title":"Effects of anthropomorphism and accountability on trust in human robot interaction,","cited_arxiv_id":null,"evidence_quote":"Shows that anthropomorphic appearance increases trust in human-robot interaction, motivating the study's focus on appearance."},{"cited_title":"The effects of overall robot shape on the emotions invoked in users and the perceived personalities of robot,","cited_arxiv_id":null,"evidence_quote":"Demonstrates that the shape of individual robot components influences perceived personality, supporting feature-level analysis of morphology."},{"cited_title":"A database for kitchen objects: Investigating danger perception in the context of human-robot interaction,","cited_arxiv_id":null,"evidence_quote":"Offers evidence that kitchen objects are perceived as more dangerous when handled by a robot, framing safety-related placement decisions."},{"cited_title":"A benchmark for recipe understanding in artificial agents,","cited_arxiv_id":null,"evidence_quote":"Supplies the two recipes whose tasks were used in the mise en place scenario."},{"cited_title":"Virtual field studies: Conducting studies on public displays in virtual reality,","cited_arxiv_id":null,"evidence_quote":"Supports the use of virtual reality as a setting where behavior resembles real-world interactions, justifying the VR methodology."},{"cited_title":"General attitudes towards robots scale (GAToRS): A new instrument for social surveys,","cited_arxiv_id":null,"evidence_quote":"Measures participants' general attitudes toward robots, used as a baseline in the participant sample."}],"review_version":1}