{"id":"2df45ce1-6a77-4169-bea3-da002e0f40c9","arxiv_id":"2601.20466","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM-based planetarium pilots are unreliable replacements for human pilots but show promise as co-pilots for reducing workload and multitasking in live shows.","lead":"Researchers built an AI assistant that listens to planetarium guides and controls OpenSpace's cameras, time, and visuals by voice. Seven professional guides tested it and agreed it cannot yet replace a human pilot but may be useful as a co-pilot to reduce workload and enable multitasking.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The experiment evaluates AI as a standalone pilot, never the proposed human+AI co-pilot configuration, so the workload-reduction claim is extrapolated, not measured.","rationale":"The paper is an honest pilot study, and the authors explicitly frame the co-pilot idea as a direction. However, the abstract's central claim states the AI 'could become useful as co-pilots to reduce workload of human pilots and allow multitasking.' For this claim to hold, the evidence would need to show that adding an AI to the human-pilot workflow reduces measured workload or enables multitasking. The study never instantiates that workflow: participants used the AI as the only pilot (reactive or proactive), and the human baseline is described from interviews rather than observed under the same protocol (Section 4). The only direct evidence for asset-toggling help is P3's anecdote about a hypothetical future ability (Section 5.3). Given that the system required repeated interventions and one participant could not run it at all (Section 4.2), the conditions under which the qualitative impressions were gathered were compromised. Thus the co-pilot workload-reduction claim is not supported by the current empirical design; it is a reasonable hypothesis for future work. The proposed test—a three-condition comparison with the actual co-pilot condition included—would directly adjudicate the claim. This aligns with the reader's CONDITIONAL verdict, but our emphasis on the missing co-pilot condition is more specific than the reader's baseline-comparison concern.","tokens_in":8411,"tokens_out":5843,"duration_ms":65435,"concrete_test":"Run a within-subjects study with at least 8 professional guides in the same dome, comparing three conditions: (1) human pilot alone (operating OpenSpace with the existing UI), (2) AI pilot alone (the paper's reactive or proactive mode), and (3) a true co-pilot condition where a human pilot and the AI operate together, with the AI handling asset toggling while the human controls camera. Use the same 10-minute simulated tour script for all conditions, counterbalanced. Measure NASA-TLX workload, number of successful asset toggles, camera-control errors (e.g., dark-side zoom-ins), and count of human interventions. If condition (3) does not show significantly lower workload or improved performance relative to condition (1), the central co-pilot claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that AI could serve as co-pilot to reduce human pilot workload and enable multitasking—is not directly tested by the study's design. The experiment compared two AI-only modes (reactive vs proactive), with five guides acting as sole pilots via voice; there was no condition in which a human pilot worked alongside the AI. The human-pilot baseline was instead reconstructed from interviews with two experts (Section 4), and the co-pilot benefit is supported only by participants' speculative statements (e.g., P3's 'It would be great if AI could add all of them at once,' Section 5.3). Moreover, the system's instability (Section 4.2: first author had to intervene several times; P4 could not run) undermines the reliability of the AI-only conditions, and the paper states that quantitative log analysis is deferred ('At this stage, we focus on reporting on the insights from the interviews,' Section 4.3). Therefore, the claim is an extrapolation from qualitative impressions under faulty conditions, not an empirical result. The workload-reduction conclusion is not internally inconsistent, but it is unsupported by the present evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes the design and evaluation of an LLM-based conversational pilot for the OpenSpace planetarium visualization system, intended to assist guides during live dome shows. The system operates in two modes: a reactive mode in which the guide triggers voice commands with a handheld microphone, and a proactive mode in which the system continuously listens and executes implicit commands. Five professional guides from the authors' institution conducted simulated shortened shows in a real dome, and two additional experts provided interview-based descriptions of human piloting practice. The paper reports qualitative thematic analysis and an error-case taxonomy. The central claim is that, while AI pilots are not yet able to replace human pilots, they could serve as co-pilots to reduce human pilot workload and enable multitasking.","tokens_in":8632,"tokens_out":2875,"duration_ms":36321,"significance":"If the central claim were empirically supported, the work would be a useful early contribution to an emerging area: LLM-driven interaction with visualization in high-stakes, time-critical, public-facing settings. The paper's strengths are its concrete system implementation, the comparison of reactive vs. proactive conversational modes, the honest reporting of failures and mixed participant preferences, and the articulation of an error taxonomy with concrete failure examples. The paper also identifies a genuinely under-explored design direction, interaction recommender systems, that could inform future work. However, the significance is currently limited because the headline conclusion about workload reduction is an extrapolation from subjective impressions and speculative participant statements rather than a measured outcome of the study's design.","major_comments":[{"comment":"The central claim that AI can act as a co-pilot to reduce human workload and enable multitasking is not directly tested. The study compares two AI-only conditions (reactive and proactive), in which the guide is the sole pilot and the AI executes commands; there is no condition in which a human pilot works alongside the AI. The workload-reduction conclusion rests on participants' speculative statements (e.g., P3's satellite anecdote in §5.3) and on a human-pilot baseline reconstructed from two expert interviews, not on any measured comparison. RQ1 is phrased as 'AI co-pilots compared to human co-pilots,' but the experiment never instantiates the co-pilot configuration. To support the claim, the paper needs either a condition with human+AI teaming, or a substantial reframing of the conclusion as a hypothesis for future work rather than a study result.","section":"§5.3 and overall design (RQ1)"},{"comment":"The evidence base for comparing AI piloting with human piloting is fragile. The human-pilot baseline is based on interviews with two consulting experts (C1, C2) rather than on logged or observed human pilot performance. In the AI conditions, the first author had to intervene several times because of software instability, and P4 could not run the system at all, instead watching P5's session. The paper states in §4.3 that 'at this stage, we focus on reporting on the insights from the interviews' and defers analysis of the collected logs. Consequently, the reported comparisons are qualitative impressions from a small, partly interrupted sample, and the abstract's 'results show' phrasing overstates the evidentiary strength. The paper should clearly label the workload-reduction and co-pilot statements as preliminary hypotheses and temper the abstract accordingly.","section":"§4.2 and §4.3"},{"comment":"The error-case analysis is presented as a substantive result, but it is anecdotal. Section 4.3 says the system logs contain latency, number, and success rate of interventions, yet §4.3 also says the paper focuses on interview insights, and §5.4 explicitly states that a 'qualitative analysis of error cases (both their characteristics, and prevalence) is needed.' No counts, success rates, or inter-rater coding are reported. The four error dimensions are plausible and useful as a taxonomy, but they are not derived from a systematic coding of the logs. Since the paper claims an 'error-case analysis from the system operation logs' in §4, this inconsistency should be resolved—either by reporting quantitative error data or by explicitly labeling the taxonomy as a set of observed examples rather than an analysis.","section":"§5.4 and §4.3"}],"minor_comments":[{"comment":"The abstract says '7 professional guides,' but §4.1 describes five expert participants (P1–P5) and two consulting experts (C1, C2). The two experts did not use the system. Please clarify the wording to distinguish study participants from expert interviewees.","section":"Abstract and §4.1"},{"comment":"P4 could not run the system and instead observed P5; the subsequent joint interview means the effective number of independent system-use sessions is four, not five. This should be stated explicitly in the participants and protocol description.","section":"§4.2"},{"comment":"The sentence 'A qualitative analysis of error cases (both their characteristics, and prevalence) is needed' is internally inconsistent: prevalence is a quantitative measure. Moreover, 'is needed' suggests the analysis has not been performed, which conflicts with the earlier claim of an 'error-case analysis from the system operation logs' in §4. Please revise to avoid this contradiction.","section":"§5.4"},{"comment":"The ACM reference format section contains placeholder dates ('February 2018', 'Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009') and the author list has inconsistent spacing in the second author's name ('M UJTABA FADHIL JA W AD'). These formatting issues should be corrected.","section":"Title page and references"},{"comment":"The finding that the reactive mode 'added to the mental load' is reported without a direct comparison of measured cognitive load. Since cognitive load is a central concept in the paper's argument, it would help to define how the authors infer cognitive load from the interviews and to acknowledge that this is a perceived, not measured, construct.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper reports an interesting exploratory study with an honestly described set of limitations. The gap between the study design and the central claim—workload reduction from a human+AI co-pilot configuration that was never tested—is the main reason for major revision. I would encourage the authors to either add such a condition in a follow-up or, more realistically, rewrite the abstract and conclusion to present the co-pilot claim as a design-level hypothesis grounded in qualitative feedback, not as an empirical result. The paper is within scope for a CSCW/HCI venue, but the evidentiary bar needs to match the claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper is a cleanly written, honest pilot of an LLM-driven planetarium pilot in OpenSpace, comparing reactive versus proactive voice control with five expert guides. That part is worth reading. The second thing is that the abstract's claim—AI could serve as a co-pilot to reduce workload—is not actually tested. The study compared two AI-only modes; there was no human-plus-AI condition. The co-pilot benefit comes from participant speculation and from a human baseline reconstructed through interviews with two experts, not from observing a human and an AI working together.\n\nWhat is genuinely new: applying proactive LLM agents to navigation in a live full-dome setting, and the qualitative error taxonomy in Section 5.4 (detection, reasoning, context, naturalness) is a useful lens for evaluating multimodal interaction agents. The authors are candid about failures: software instability, the P4 drop-out, mixed preferences between modes. That honesty counts.\n\nThe soft spots are real but not fatal to the exploratory framing. Five participants from the authors' own institution, a simulated show with the first author intervening repeatedly, no quantitative log analysis—all acknowledged in Section 4.3. The more serious issue is the conceptual gap between what they measured (AI alone) and what they conclude (AI as co-pilot reduces workload). The stress-test note gets this right. P3's remark about adding satellite assets is suggestive, and the 'interaction recommender system' direction is plausible, but it is a research direction, not an empirical result.\n\nNo circularity problem; conclusions are inductive from participant feedback, and the self-citations establish lineage rather than determining the outcome. The paper does not oversell—it explicitly says AI will not replace human pilots.\n\nThis is for researchers working on LLM-driven visualization, science-center tech staff, and anyone designing proactive assistants in public-facing settings. It deserves a serious referee: the question is current, the implementation is described well enough to reproduce, and the qualitative evidence is usable in a workshop paper or as a preliminary study. I would not accept it as-is. I would send it to review with instructions to address the co-pilot extrapolation, either by adding a small co-pilot condition or by softening the abstract's workload claim to match the evidence.","headline":"An honest little pilot study with a useful error taxonomy, but the headline workload-reduction claim rests on an extrapolation, not on the co-pilot condition the paper actually proposes.","tokens_in":9099,"tokens_out":3036,"would_cite":false,"duration_ms":34534,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM planetarium pilots can cut human workload as co-pilots, but cannot replace pilots.","keywords":["LLM","Visualization","Public Spaces","Conversational AI","Proactive AI","Planetarium","Human-AI collaboration","OpenSpace"],"falsifier":"Run real public planetarium shows with the same AI co-pilot and measure human pilot workload and experimenter interventions; if the AI requires intervention nearly as often as it saves actions, or if guides report no workload drop, the central claim fails. A sharper version would compare AI co-pilot shows head-to-head with human-human pilot pairs on the same measures.","tokens_in":8323,"feed_emoji":"🌌","tokens_out":5034,"duration_ms":56938,"temperature":0.7,"pith_summary":"The paper argues that an LLM-based conversational agent can listen to a planetarium guide's speech and execute visualization commands—moving the camera, changing simulation time, and toggling assets—but that this AI pilot lacks the timing, camera control, and audience-reading skills of a human pilot. The central claim is that instead of replacing the human pilot, the AI is best used as a co-pilot: it can handle asset preparation and toggling while the human keeps control of pacing and camera motion, reducing workload. A comparative study with five guides testing the system and two experts providing the human-pilot baseline found mixed preferences between reactive and proactive modes, with proactive feeling more fluid but less reliable. The paper concludes that AI pilots are not a replacement and likely never will be, but are useful for multitasking, preparation, and onboarding novice pilots.","feed_headline":"LLM planetarium pilots work best as co-pilots, not replacements","feed_subtitle":"An AI pilot in planetarium software toggles assets and prepares actions; humans keep camera and pacing.","key_machinery":"The carrying mechanism is the conversational AI pilot built on the astrophysics visualization software OpenSpace and a low-latency multimodal language model. It receives the guide's speech (in reactive mode through a hand-triggered microphone, in proactive mode through streaming speech-to-text), interprets the speech as commands, and dispatches tool calls through OpenSpace's Lua API to travel between scene nodes, toggle asset visibility, change simulation time, or do nothing via a 'no-op' call. The comparison between reactive and proactive modes is the experimental pivot.","core_discovery":"The paper's central claim is that an LLM-based pilot, which listens to a guide's speech and executes commands like moving the camera, changing simulation time, and toggling assets, is not able to replace a human pilot in a live planetarium show. The decisive observation is that the AI is good at toggling assets and preparing actions but poor at camera control and pacing, which demand temporal awareness and precision. The paper therefore proposes a division of labor: keep a human pilot for camera and pacing, and let the AI act as a co-pilot to reduce cognitive load and enable multitasking. Five guides tested reactive and proactive modes; proactive felt more natural but less reliable, while re","pith_inferences":["If the co-pilot division of labor holds, the same pattern could transfer to other live visualization settings, such as surgery or forensic reconstruction, where a human controls continuous motion while AI preloads context.","The paper's 'interaction recommender system' idea suggests a less autonomous intermediate step: the AI suggests actions rather than executing them, which may be more robust before streaming speech and latency improve.","The error dimension spectrum could be formalized into a benchmark with human-coded ground truth, making future AI pilot comparisons quantitative rather than interview-based.","A reading of the mixed proactive/reactive preference: reliability is the gating factor; if proactive mode's reliability approaches reactive mode's, its fluidity advantage may make it the default."],"forward_implications":["Planetarium shows would keep a human pilot for camera control and pacing, while the AI co-pilot handles toggling and preparing visual assets.","Proactive listening, once reliability improves, can reduce mental load and preserve narrative flow because guides no longer need to press a trigger.","The error dimensions identified—detection, reasoning, context, and naturalness—offer a concrete checklist for building and evaluating AI pilots.","AI could assist during show preparation and onboarding, letting novice pilots assemble and test storytelling blocks in a low-risk setting."],"fun_headline_variants":["LLM planetarium pilot: co-pilot, not replacement","AI handles assets, humans maintain camera in planetarium shows","Planetarium AI pilot lacks timing, wins as co-pilot","Seven guides agree: LLM pilots need a human at the camera"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The conclusion that AI co-pilots reduce workload rests on the assumption that the experience of five guides in simulated shows—where the experimenter had to intervene repeatedly because the software was unstable—and the expectations of two interviewed experts predict what would happen in real public shows.","fun_headline_variants_meta":{"raw":{"variants":["LLM planetarium pilot: co-pilot, not replacement","AI handles assets, humans maintain camera in planetarium shows","Planetarium AI pilot lacks timing, wins as co-pilot","Seven guides agree: LLM pilots need a human at the camera"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000263,"raw_usage":{"total_tokens":1395,"prompt_tokens":664,"completion_tokens":731,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":659}},"tokens_in":408,"tokens_out":731,"duration_ms":7530,"temperature":1.0,"reasoning_tokens":659,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T07:19:13.070169+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run real public planetarium shows with the same AI co-pilot and measure human pilot workload and experimenter interventions; if the AI requires intervention nearly as often as it saves actions, or if guides report no workload drop, the central claim fails. A sharper version would compare AI co-pilot shows head-to-head with human-human pilot pairs on the same measures.","supporting_citations":[],"review_version":1}