{"id":"7bb72235-8b28-42a7-8a3b-a2e951663f39","arxiv_id":"2607.01044","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Adding a communication module to social navigation models improves episode success by 10 percentage points in simulated multi-human environments and remains robust to natural language inputs.","lead":"This paper introduces CommNav, where robots proactively ask humans for information to locate targets in multi-person settings instead of only avoiding collisions. A smart generalist might read it to see how adding simple communication changes robot performance in social environments like homes or care facilities.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Habitat 3.0c information-exchange protocols lack validation against real human responses, making simulator-only gains in Episode Success non-transferable.","rationale":"The reader's weakest assumption directly identifies the load-bearing condition for both headline claims. Because the full manuscript remains simulation-only (no physical-robot or real-human interaction results are described), the concern stands and the UNVERDICTED verdict with low confidence is unchanged.","tokens_in":1746,"tokens_out":315,"duration_ms":19762,"concrete_test":"Deploy the trained COMM policy on a physical robot in a controlled multi-resident room; collect live human answers to the same query templates used in the human study, run 50 episodes each with structured vs. natural-language inputs, and compare success rates to the Habitat 3.0c numbers. A >5pp drop in the natural-language vs. structured gap falsifies transfer.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The 10pp Episode Success improvement and the statistical equivalence between perfect structured data and colloquial natural language both rest on interactions generated inside the extended Habitat 3.0c simulator. The paper defines multi-human environments and information-exchange protocols, yet provides no comparison of simulated responses (e.g., sighting reports, location answers) to actual human behavior collected in physical settings. If real humans produce more variable, ambiguous, or off-topic replies than the simulator's protocol, the COMM module's reported robustness and performance delta become artifacts of the virtual environment rather than intrinsic properties of the policy.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces CommNav, a new task for robots to proactively gather information via human-robot communication to locate targets in multi-human environments. It extends Habitat 3.0 into Habitat 3.0c with multi-human support and information-exchange protocols, adds a COMM module to a state-of-the-art social navigation baseline, and reports a 10 percentage-point gain in Episode Success. Experiments further show that the policy remains robust when switching from perfect structured data to LLM-generated or human-collected colloquial natural-language instructions.","tokens_in":1861,"tokens_out":420,"duration_ms":21742,"significance":"If the empirical results hold under proper statistical reporting, the work provides a concrete demonstration that explicit communication can substantially improve multi-person social navigation performance. The new simulator variant, the communication pretext pre-training, and the inclusion of a human study for language robustness constitute clear contributions to the field of assistive robotics.","major_comments":[{"comment":"Abstract: the central claims of a '10 percentage-point improvement in Episode Success' and 'episode success statistically similar' to the perfect-structured-data baseline are stated without error bars, number of episodes or runs, baseline model name, or any statistical test results, preventing assessment of whether the reported delta is reliable or significant.","section":"Abstract"},{"comment":"Abstract (paragraph on Habitat 3.0c creation): the information-exchange protocols (sighting reports, location answers, etc.) are defined inside the simulator but receive no validation against real human responses collected in physical settings; if real humans produce more variable or off-topic replies, both the 10 pp gain and the claimed robustness to colloquial language rest on untested simulator assumptions.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract would be clearer if it named the specific state-of-the-art social navigation model used as the baseline for the COMM ablation.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address each major comment below and indicate the revisions we will make to strengthen the manuscript.","responses":[{"response":"We agree that these statistical details are necessary for proper assessment. In the revised abstract and main text, we will specify the number of evaluation episodes (500 per condition), report standard errors from 5 independent runs, name the baseline model explicitly, and include the results of the statistical tests (paired t-tests) used to support both the 10 percentage-point improvement and the claim of statistical similarity.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claims of a '10 percentage-point improvement in Episode Success' and 'episode success statistically similar' to the perfect-structured-data baseline are stated without error bars, number of episodes or runs, baseline model name, or any statistical test results, preventing assessment of whether the reported delta is reliable or significant."},{"response":"The protocols model plausible information exchanges to enable controlled study of the COMM module. Our human study already collects real colloquial instructions to test language robustness, which partially addresses variability in human responses. We will revise the manuscript to explicitly state the simulation assumptions, note that full physical validation of protocol dynamics lies outside the current scope, and discuss this as a limitation.","revision_made":"partial","referee_comment":"[Abstract] Abstract (paragraph on Habitat 3.0c creation): the information-exchange protocols (sighting reports, location answers, etc.) are defined inside the simulator but receive no validation against real human responses collected in physical settings; if real humans produce more variable or off-topic replies, both the 10 pp gain and the claimed robustness to colloquial language rest on untested simulator assumptions."}],"tokens_in":1362,"tokens_out":387,"duration_ms":29734,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper's main takeaway is that adding their COMM module to a social navigation baseline lifts episode success by 10 points in simulation, and the policy stays roughly as effective when the input switches from perfect structured data to colloquial natural language drawn from a human study.\n\nWhat is new is the CommNav task itself, where the robot actively queries residents for sighting and location information, plus the dedicated communication module and the specific robustness result to natural language. They also built Habitat 3.0c to handle multi-human environments and information exchange.\n\nThe work does a clean job showing that explicit communication helps in multi-person navigation and that pre-training on a communication pretext task handles sparse signals. The human study for generating colloquial instructions is a reasonable step toward realism on the input side.\n\nThe soft spot is the simulator. All reported gains and the equivalence between structured and natural language rest on the custom information-exchange protocols inside Habitat 3.0c. The abstract gives no evidence that those simulated responses match how real people actually reply in comparable situations, so the transfer story remains open. The abstract also omits error bars, exact baseline details, and statistical tests, which leaves the central claim harder to assess without the full methods.\n\nThis is for people working on social and assistive robotics. It introduces a concrete new task with some empirical backing, so it deserves a serious referee to sort out the sim validation and methods details.","headline":"The paper adds a CommNav task and COMM module that delivers a 10pp episode success gain in Habitat 3.0c plus robustness to colloquial language, but the gains sit on unvalidated simulator response protocols.","tokens_in":2339,"tokens_out":374,"would_cite":false,"duration_ms":18025,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Robots improve navigation success by 10 points when they ask humans for directions in crowds.","keywords":["social navigation","human-robot communication","multi-agent environments","assistive robots","Habitat simulator","CommNav task"],"falsifier":"A physical experiment with real robots and human participants in a multi-person setting that measures whether episode success rises by a similar margin when the communication module is added versus when it is absent.","tokens_in":2639,"feed_emoji":"🤖","tokens_out":587,"duration_ms":26064,"temperature":0.7,"pith_summary":"The paper introduces CommNav, a task in which robots must locate specific people among multiple residents by actively requesting information through communication rather than relying only on reactive avoidance. It extends the Habitat simulator to Habitat 3.0c to support multi-human environments and information-exchange protocols. Adding a communication module (COMM) to an existing social navigation model raises episode success by 10 percentage points. The resulting policy maintains statistically similar performance when humans reply in colloquial natural language instead of perfect structured data.","feed_headline":"Robots gain 10 points navigation success by asking humans","feed_subtitle":"A new communication module lets robots query residents about target locations, matching structured-data performance even with casual speech.","key_machinery":"The COMM communication module that lets the robot issue queries about sightings and locations and integrates the resulting responses into the navigation policy.","core_discovery":"In CommNav, robotic agents seek assistance from residents by requesting details about recent sightings, locations, and movements of target individuals. The addition of the COMM module to a state-of-the-art social navigation model produces a 10 percentage-point gain in Episode Success. Pre-training COMM on a communication pretext task addresses infrequent interaction signals, and the navigation policy remains robust to natural colloquial human language, reaching episode success rates statistically similar to those obtained with perfect structured data.","pith_inferences":["The same query-and-response pattern could reduce reliance on complete prior maps in highly dynamic indoor spaces.","Robustness to casual language suggests the approach may work with untrained bystanders without requiring special phrasing.","If simulator results transfer, similar communication modules could be tested in delivery or search tasks that also require locating moving targets among people."],"forward_implications":["Explicit human-robot communication substantially raises multi-person navigation performance.","Pre-training on a communication pretext task improves handling of occasional interaction signals.","Navigation policies trained on LLM-generated or human-collected colloquial instructions perform comparably to those using perfect structured data."],"fun_headline_variants":["Robots gain 10 points asking humans about targets","Communication module gives 10-point success gain","Robots perform similarly with casual human language","COMM pre-training manages infrequent interaction signals"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The extended Habitat 3.0c simulator and its information-exchange protocols produce interaction patterns that transfer to real human-robot communication in physical environments.","fun_headline_variants_meta":{"raw":{"variants":["Robots gain 10 points asking humans about targets","Communication module gives 10-point success gain","Robots perform similarly with casual human language","COMM pre-training manages infrequent interaction signals"]},"model":"grok-4.3","cost_usd":0.01017,"raw_usage":{"total_tokens":4516,"prompt_tokens":681,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":101699500,"prompt_tokens_details":{"text_tokens":681,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3781,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":681,"tokens_out":54,"duration_ms":30354,"temperature":1.0,"reasoning_tokens":3781,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-02T11:11:32.141942+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A physical experiment with real robots and human participants in a multi-person setting that measures whether episode success rises by a similar margin when the communication module is added versus when it is absent.","supporting_citations":[],"review_version":1}