{"id":"8f0cafe5-035a-470f-92df-30796f3aadec","arxiv_id":"2607.26423","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A mixed-reality 3D sandtable lets single operators supervise fleets of 5–15 simulated inspection drones, with situational awareness degrading beyond roughly 10 drones.","lead":"FleetScape is a mixed-reality sandtable for supervising drone fleets that shows drone positions, mission progress, and safety alerts in one 3D space. A six-pilot study found the interface supports awareness up to about 10 drones, above which operators switch to reactive exception-based monitoring.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SAGAT scoring in §B.4 (Jaccard on sets) is incompatible with 'How many…' queries in Table 3, making the SA-decline evidence for the ~10-drone limit unverified.","rationale":"The reader's weakest_assumption is external validity: that the simulated building-inspection mission may not transfer to real fleet operations. That is a legitimate concern, but it is explicitly conceded in the Limitations and does not threaten the internal validity of the in-simulation results. The more immediate threat is internal: the SAGAT numbers in Figure 4 are the paper's clearest quantitative evidence for situational-awareness decline, yet the scoring formula in Appendix B.4 presupposes set-valued answers to queries that Table 3 often phrases as count questions. This is not a matter of intent or statistical power; it is an unaddressed specification gap. Without raw data or a clarification, the reader cannot verify the central quantitative trend. The recommendation remains CONDITIONAL rather than REJECT or UNVERDICTED because the paper also provides converging qualitative evidence, subjective SART/Bedford ratings, and a fan-out calculation, so the central claim is not wholly dependent on SAGAT. However, the acceptance condition should explicitly require the authors to release the SAGAT response format and a re-analysis using the correct scoring metric for count-based questions; if the decline does not survive that re-scoring, the empirical limit is unsupported.","tokens_in":19244,"tokens_out":9384,"duration_ms":99038,"concrete_test":"Obtain the raw SAGAT response logs or answer sheets. For every \"How many...\" item, determine whether the participant gave a numeral or an enumerated set. If responses are numerals, re-score those items with a count-based metric (e.g., 1 − |G−R|/max(G,R), or exact-match accuracy) and recompute Figure 4A/B. If the drone-status and mission-comprehension declines from 5→10→15 persist under this corrected scoring, the concern is resolved as a reporting omission; if they weaken or reverse, the claimed ~10-drone limit is not supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing concern is an internal inconsistency between the SAGAT scoring protocol and the SAGAT questions. Appendix B.4 states that perception/comprehension queries are scored with the Jaccard index on a ground-truth set G and response set R (Eq. 4). But Table 3 contains multiple count queries, e.g., \"How many buildings have finished surface(s)?\", \"How many drone(s) is in one of the stoppage statuses...\", and \"how many drone(s) are likely to stop...\". If participants answered these with numerals, the Jaccard formula is undefined, and the reported drone-status SA decline (0.86→0.63→0.61, Fig. 4B) and mission-comprehension decline (0.85→0.58) cannot be computed from the stated method. If instead participants were asked to enumerate sets, that protocol is not reported. Since these scores are the primary quantitative support for the '~10 drones' boundary, the empirical limit is insecure until the raw response format is clarified.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents FleetScape, a mixed-reality sandtable interface for supervising fleets of autonomous drones in a building-inspection task. The system combines a 3D sandtable with a minimap, status board, and per-drone control panel to support spatial reasoning and transitions among manual control, semi-autonomous route planning, and autonomous inspection. The authors report an exploratory user study with six experienced drone pilots operating fleets of 5, 10, and 15 drones in a Unity-based simulator. They find that mission perception improves with fleet size while mission comprehension and drone-status situational awareness decline, and they propose a fan-out-derived estimate of about 10 drones as the practical upper bound for single-operator spatial supervision. The paper also derives qualitative design implications around adaptive visualization, control grouping, and pre-mission planning.","tokens_in":19508,"tokens_out":4425,"duration_ms":47066,"significance":"If the results are reliable, the paper makes a useful design contribution by framing drone-fleet supervision as a spatial-interaction problem and by providing a concrete scaling boundary for one operator. The use of professional pilots, the transparently described simulation, the SAGAT freezes grounded in simulator ground truth, and the explicit fan-out calculation are strengths. The manuscript also candidly acknowledges its exploratory nature and the simulated setting. However, the central empirical claim — the decline in situational awareness and the 'around 10 drones' limit — rests on the SAGAT scores, and the scoring protocol as written does not fit all of the reported questions. The study is also too small and too underpowered to support strong quantitative claims without inferential statistics. The design implications remain valuable even if the precise number is treated as provisional.","major_comments":[{"comment":"The scoring protocol in Eq. (4) applies the Jaccard index to a ground-truth set G and a response set R for perception and comprehension questions. However, Table 3 contains count questions such as 'How many buildings have finished surface(s)?', 'How many drone(s) is in one of the stoppage statuses...', and 'how many drone(s) are likely to stop the automation within the next 1 minute?'. If participants answered these with a numeral, Eq. (4) is undefined. If they were instead asked to enumerate the set of buildings/drones, that response format is not reported. Since the drone-status SA decline (0.86→0.63→0.61) and mission-comprehension decline (0.85→0.58) in Fig. 4 are the primary quantitative support for the ~10-drone limit, please specify the response format per question and the exact scoring rule used. If counts were converted to sets, describe the conversion and justify it.","section":"Appendix B.4 / Table 3"},{"comment":"With n=6 and the reported standard deviations being large (e.g., mission comprehension SDs of 0.28 and 0.26; drone comprehension SD of 0.33 in the 15-drone condition), the directional statements 'decreased', 'dropped', and 'a limit was observed' are not backed by inferential statistics. The paper itself labels the study exploratory, yet the abstract and conclusion present 'around 10 drones' as an empirical finding. Please either provide appropriate statistical tests or effect sizes with a clear caveat about the small sample, or rephrase the central scaling claim as a descriptive observation to be confirmed in larger studies.","section":"§4.2.1 / Fig. 4"},{"comment":"The fan-out estimate of FO≈10.6 is computed from interaction and activity times measured in the same prototype and simulation, and it is used to select the 5/10/15 fleet-size conditions. The Discussion and Limitations then invoke this estimate as supporting the observed 5–10 range. This is not a fully independent validation. Please clarify that the SAGAT-based SA decline is the primary evidence for the scaling limit and that the fan-out calculation is a design heuristic. Also clarify whether the timing data in Appendix C.1 come from the trial-stage pilots or the final participants, since the current text is ambiguous.","section":"§4.1.2 / Appendix C.1 / §6"}],"minor_comments":[{"comment":"The text refers to 'Per-drone performance (Figure 4.E)', but the Figure 4 caption lists per-drone productivity as panel (C), with Bedford workload as (D) and SART as (E). Please correct the cross-reference.","section":"§4.2.1"},{"comment":"The line 'IT negligible: SA takeoff≈5s' is unclear. If takeoff time is negligible, why is 5 s listed? Please clarify the notation and the units for the rates (s⁻¹).","section":"Appendix C.1"},{"comment":"The projection-question scoring is described as binary and based on 'coherence between the answer and the reason given'. This is underspecified. Please define what counts as coherent, especially for the mission-projection question ('which building will be finished next?').","section":"Appendix B.4"}],"recommendation":"major_revision","confidential_remarks":"The SAGAT scoring inconsistency is the main technical concern. It is potentially fixable with a clear description of the response format and scoring rule, so I do not recommend outright rejection. The paper would also be strengthened by adding inferential statistics or softening the quantitative claims. The design discussion and qualitative findings are genuinely useful for the UIST community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know one thing up front: FleetScape is a legitimate design exploration, not just a tech demo. An MR sandtable with layered mission/safety data and smooth control transitions is a sensible spatial-supervision design, and they tested it with six professional pilots at 5/10/15 drones. The main empirical claim—situational awareness degrades somewhere around 5–10 drones and operators shift to reactive, exception-based supervision—is plausible and, importantly, qualitatively well-supported by the interview quotes. The design-space table and the interface description are detailed enough to build on.\n\nWhat is genuinely new: this is the first system I know of that combines a 3D sandtable, layered environmental/mission/safety visualization, and per-drone control transitions in a single spatial context, and the scaling evaluation across 5–15 drones is a real step beyond single-drone MR interfaces. The fan-out computation in Appendix C.1 is transparent, and the Limitations section is unusually candid: they explicitly say the fan-out estimate of 10.6 is an upper bound and that real-world latency and risk would lower it.\n\nNow the soft spots, in proportion.\n\nFirst, the stress-test concern is valid and it lands on a load-bearing part of the quantitative story. Appendix B.4 says perception and comprehension queries are scored with the Jaccard index on sets. But Table 3 includes at least three count questions—\"How many buildings have finished surface(s)?\", \"How many drone(s) is in one of the stoppage statuses...\", and \"How many drone(s) are likely to stop...\". A numeral does not have a Jaccard index. If participants instead enumerated sets, that protocol is not reported. As written, the reported drone-perception decline (0.88→0.59→0.63) and mission-comprehension decline (0.85→0.58) cannot be reproduced from the stated method. This doesn't kill the paper's thesis, since the qualitative data and the non-count SAGAT items point the same way, but the \"around 10 drones\" boundary needs that clarified.\n\nSecond, the standard exploratory weaknesses: n=6, high standard deviations, no baseline or comparison condition, and a simulator that abstracts away latency and risk. The paper acknowledges all of this, so I won't belabor it. The fan-out interaction times were measured on the same prototype used to choose the 5/10/15 conditions, which is mildly circular, but the authors flag it.\n\nThird, the trial stage included amateur pilots, and the final stage included one amateur; that is minor and the demographics are clearly reported.\n\nOn balance, the reader's CONDITIONAL verdict is fair. My own take is slightly more positive on the design contribution and slightly more cautious on the quantitative boundary claim until the scoring protocol is fixed.\n\nRecommendation: yes, send this to peer review. The system is buildable, the qualitative findings are credible, and the design implications are actionable. A serious referee should ask the authors to restate the response format for each SAGAT item and re-report the scores accordingly. If the numbers change, the 10-drone boundary might soften, but the paper would still contribute.","headline":"Solid exploratory system paper with a real design contribution, but the SAGAT scoring appendix has an internal ambiguity that makes some reported SA numbers unverifiable; the qualitative strategy-shift finding holds up.","tokens_in":19983,"tokens_out":2418,"would_cite":true,"duration_ms":31205,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reframing drone fleet supervision as spatial interaction on a mixed-reality sandtable supports situational awareness up to about 10 drones, beyond which awareness degrades and operators fall back to reactive monitoring.","keywords":["Mixed Reality","Spatial Interfaces","Human-Drone Interaction","Multi-drone supervision","Situational awareness","Sandtable","Swarm robotics","Fan-out"],"falsifier":"A field test with real drones, real latency, and real GPS/wind conditions—keeping the same autonomy and interface—would settle the claim: if operators maintain drone-level awareness past 10 drones, the limit is a simulation artifact; if awareness collapses earlier, the ceiling is lower. Measuring true interaction times in the field and recomputing the fan-out ratio would directly test the 10.6 estimate.","tokens_in":19134,"feed_emoji":"🚁","tokens_out":6505,"duration_ms":62268,"temperature":0.7,"pith_summary":"The paper argues that supervising a drone fleet is fundamentally a spatial task, and that a 3D mixed-reality sandtable—a miniature of the operational environment that layers mission, safety, and environmental data in one place—can support an operator's situational awareness better than scaled-up single-drone interfaces. To test this, the authors built a building-inspection simulation and ran a study with six professional pilots managing fleets of 5, 10, and 15 drones. They found that pilots could maintain awareness and effective control up to roughly 10 drones; at 15 drones, awareness and per-drone productivity dropped, and pilots shifted from proactive optimization to reactive, exception-based monitoring. If the finding transfers to the field, it sets a practical scaling boundary for single-operator drone fleet supervision and motivates design changes—adaptive visualizations, grouping, pre-mission planning—beyond that boundary.","feed_headline":"10 drones is the single-operator fleet ceiling","feed_subtitle":"Mixed-reality sandtable keeps fleet oversight clear up to ~10 drones; beyond that, situational awareness drops.","key_machinery":"The central mechanism is the spatially grounded sandtable: a 3D miniature of the mission environment into which all data layers are projected, so the operator perceives relationships between drones, tasks, and constraints directly instead of through mental rotation. Supporting it are focus-aware drone miniatures (detailed for the selected drone, simplified for the rest), layered visualizations (building boundaries, waypoints, point clouds), and a control-transition design that keeps mode state visually consistent across the sandtable, minimap, and status board. The fan-out metric—activity time divided by interaction time—serves as the predictive scaling tool, yielding the ~10.6 drone estimat","core_discovery":"FleetScape's central claim is that fleet supervision can be restructured as spatial interaction: rather than juggling separate camera feeds, telemetry panels, and 2D maps, an operator works with a unified 3D sandtable where drones, waypoints, building boundaries, point clouds, and alerts sit at their true positions. The design allows fluid shifts between global supervision, waypoint replanning, and manual control of one drone without losing mission context. In a study with six expert pilots, drone-level (safety) awareness fell from 0.86 at 5 drones to 0.61 at 15, and per-drone throughput from 39.6 to 32.1 covered waypoints, while mission-level awareness held fairly steady. The authors read t","pith_inferences":["A physical deployment with real latency, communication dropouts, and actual risk would likely push the practical ceiling below 10 drones; the paper itself describes 10.6 as an upper bound from simulation.","The observed shift from proactive optimization to reactive, exception-based monitoring may be a general signature of cognitive saturation, suggesting that scalable interfaces should be designed for exception handling (prioritization, grouping, abstraction) rather than expecting continuous spatial optimization.","The divergence between stable mission-level awareness and falling drone-level awareness implies future systems could separate the two roles: a global spatial view for survey and a separate automated attention system for individual drone safety.","A direct testable extension: repeat the same task with novice operators or with different GPS-error event rates; the ~10 boundary likely shifts with training and scenario stress, which would refine the design guidance."],"forward_implications":["A single operator can maintain situational awareness and effective control over a fleet of 5–10 drones when supervision is grounded in a spatial sandtable.","Mission-level awareness (which buildings are covered) survives fleet growth, but drone-level safety awareness declines, so interfaces must treat the two as separate design targets.","Beyond about 10 drones, per-drone inspection throughput drops and operators switch from proactive optimization to reactive, exception-based monitoring.","Scaling past this range requires adaptive visualizations, control groups, abstraction, and pre-mission planning, rather than only better real-time interaction techniques.","The fan-out estimate of approximately 10.6 gives a concrete, reusable upper bound for single-operator supervision under reliable autonomy, with 5–10 drones as the conservative practical range."],"fun_headline_variants":["FleetScape MR sandtable: clear oversight up to 10 drones","Mixed reality fleet control—awareness drops past 10","Fleet supervision via MR: 10-drone limit emerges","Spatial sandtable aids swarms, but scale caps at ~10","Drone fleet oversight in mixed reality fades beyond 10"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the simulated building-inspection mission is representative enough of real drone-fleet supervision that the observed awareness decline and the ~10 drone boundary will transfer to physical operations with real latency, communication constraints, and risk.","fun_headline_variants_meta":{"raw":{"variants":["FleetScape MR sandtable: clear oversight up to 10 drones","Mixed reality fleet control—awareness drops past 10","Fleet supervision via MR: 10-drone limit emerges","Spatial sandtable aids swarms, but scale caps at ~10","Drone fleet oversight in mixed reality fades beyond 10"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1121,"prompt_tokens":738,"completion_tokens":383,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":292}},"tokens_in":482,"tokens_out":383,"duration_ms":4946,"temperature":1.0,"reasoning_tokens":292,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T16:12:12.695417+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A field test with real drones, real latency, and real GPS/wind conditions—keeping the same autonomy and interface—would settle the claim: if operators maintain drone-level awareness past 10 drones, the limit is a simulation artifact; if awareness collapses earlier, the ceiling is lower. Measuring true interaction times in the field and recomputing the fan-out ratio would directly test the 10.6 estimate.","supporting_citations":[],"review_version":1}