{"id":"d20076a7-63cd-49f3-8c6f-1d6f0da619c9","arxiv_id":"2501.06566","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"CARIC is a new simulation benchmark for heterogeneous multi-UAV inspection, and its first competition results show no single planning strategy wins across all scenarios.","lead":"This paper introduces CARIC, a simulation benchmark for teams of camera and LiDAR drones inspecting unknown structures, and reports lessons from the first competition at IEEE CDC 2023. It analyzes the strategies and scores of the top three teams to show trade-offs between exploration detail, inspection quality, and task allocation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All three top competition teams are from the authors' own institutions, and the paper neither discloses this nor releases per-run scores; the 'lessons learned' may thus be self-assessment, weakening the central empirical claim.","rationale":"The paper's strongest claim has two components: that CARIC is a ready-to-use benchmark, and that the first competition demonstrates the absence of a robust single approach. The first component is supported by public code, documentation, and sample scenarios—credible artifacts even if the evaluation metrics are imperfect. The second component, however, lacks independent evidentiary support: all three teams analyzed in depth are from the authors' own institutions, and per-run data are withheld. The reader's weakest_assumption (metric calibration) is a modeling concern common to simulation benchmarks; it affects how well CARIC scores transfer to real inspection, but it does not by itself undermine the reported competition outcome. The author-affiliation and data-availability problem is more load-bearing because it calls into question whether the empirical demonstration is objective at all. If the competition's top performers are the benchmark designers themselves, the 'lessons learned' are a self-assessment, not an independent community result. This is not an accusation of misconduct; it is a structural conflict that must be disclosed and addressed with per-run data. The reader's CONDITIONAL verdict already includes the disclosure condition, so my assessment leaves that verdict unchanged: the paper can be accepted only after the transparency issues are resolved and the empirical claims are re-derived from independent, non-organizer teams or explicitly recast as organizer-team case studies.","tokens_in":10179,"tokens_out":10102,"duration_ms":106035,"concrete_test":"Obtain the competition records (team registration forms, per-run score sheets) from the organizers. First, verify whether any member of each top-three team is also an author of this paper or served on the CARIC organizing committee. Second, recompute the team ranking using the five per-run scores with the mean and the median, and compare against the reported max-based ranking. If all three top teams are organizer-affiliated, or if the ranking changes under mean/median scoring, then the paper's empirical conclusions must be revised to disclose the conflict and to report results for external teams separately.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that the CDC 2023 competition shows no single approach achieves robust performance across scenarios (Section VI)—rests on the achievements of the 'top three teams.' Checking author affiliations against team affiliations reveals that all three analyzed teams include members of this paper's own author team: KIOS CoE is the University of Cyprus group listed among the authors (Anastasiou, Zacharia, Papaioannou, Kolios, Panayiotou, Polycarpou); XXH is Nanyang Technological University, which includes authors Nguyen, Yuan, Xu, and Xie; STAR is Sun Yat-sen University, which includes authors Zhang and Zhou. The manuscript nowhere discloses this overlap. It also does not release per-run scores, only box plots, so an external reader cannot separate author-team performance from external teams' performance. If the competition's top results are largely produced by the benchmark designers and their close collaborators, the conclusion that CARIC exposes robust-performance gaps among the broader research community is not established. The benchmark software may still be a useful artifact, but the paper's 'lessons learned' and its claim of attracting 'innovative solutions from research teams worldwide' are not independently supported. The appropriate remedy is full disclosure of the overlap, plus release of per-run scores and a separate analysis of non-organizer teams.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CARIC, a Gazebo/RotorS-based simulation benchmark for heterogeneous multi-UAV inspection planning, in which teams of LiDAR-equipped explorers and camera-equipped photographers must inspect interest points on structures known only by bounding boxes, subject to line-of-sight communication and image-quality scoring (blur and resolution). The paper reports the CARIC competition held at CDC 2023, presents the approaches of the top three teams (KIOS CoE, XXH, STAR), and draws lessons about exploration-vs.-inspection trade-offs, inspection quality vs. detection rate, task allocation, and performance variance. The central empirical claim is that no single approach achieved robust performance across all scenarios.","tokens_in":10459,"tokens_out":2964,"duration_ms":28898,"significance":"If the benchmark is adopted by the community, CARIC would provide a useful standardized testbed for a previously underexplored problem: cooperative heterogeneous UAV inspection without a prior structural model. The paper's strengths include a publicly available, ready-to-use simulation stack with clearly defined scoring metrics (Eqs. 1–6), a concrete competition protocol with multiple runs and two hardware configurations, and an explicit statement that rankings were consistent across those configurations. The authors also provide qualitative per-scenario analyses with illustrative trajectory and detection maps. However, the empirical lessons are substantially weakened by two issues: all three analyzed teams are from the authors' own institutions, and per-run scores are not released. These concerns must be addressed before the paper's conclusions about community-wide performance can be accepted.","major_comments":[{"comment":"The three 'top teams' analyzed in Sections IV and V — KIOS CoE (University of Cyprus), XXH (Nanyang Technological University), and STAR (Sun Yat-sen University) — all include members of this paper's author list, as shown by the affiliations in the author block. The manuscript nowhere discloses this overlap. Because the lessons in Section V are drawn exclusively from these self-affiliated teams, the central claim that no single approach has achieved robust performance across scenarios (Section VI) is not an independent assessment of the broader participant pool. The paper should (a) explicitly disclose the author–team overlap, (b) release per-run scores for all teams or at least for the analyzed teams, and (c) provide a separate analysis of non-organizer teams if such data exist. Without this, the lessons learned are better characterized as a self-assessment of the organizers' own algorithms.","section":"III and V"},{"comment":"The claim that no single approach is robust across all scenarios rests on box plots for only three teams, with no statistical tests, effect sizes, or baseline comparisons. This is especially concerning because Section V-D1 reports within-team variance of over 1000 points between runs for XXH and STAR; the observed cross-scenario differences may be partly due to run-to-run noise. The paper should report per-run scores, indicate the number of runs used for each box in Figure 9, and either add significance testing or explicitly state that the cross-scenario comparison is anecdotal given n=3 teams.","section":"V-D2"},{"comment":"The resolution metric q_res uses the parameter rdes as the desired MMPP, and the introduction to Section II states that the metrics 'ensure the captured data meets standards for structural analysis.' However, no calibration or validation is provided linking q_blur and q_res to actual defect detection or structural inspection outcomes. As a result, the benchmark's rankings may not transfer to field deployment, and the claim that the scores reflect inspection quality for structural analysis is currently unsupported. Either provide a calibration study against real inspection data or soften the wording to describe these as heuristic image-quality proxies.","section":"II-D3 (Eq. 6)"}],"minor_comments":[{"comment":"There is a typo in the definition of v1: it is printed as 'v1 = f · z1/z1', which should presumably be 'v1 = f · y1/z1'.","section":"II-D2, Eq. (4)"},{"comment":"The text says teams are ranked by the max score across five runs, while Figure 9 shows box plots of 'overall scores across tests.' Please clarify whether the box plots summarize the five runs per team, and indicate the number of runs and whether any runs were excluded (e.g., due to crashes).","section":"III and Figure 9"},{"comment":"The statement 'The rankings were consistent across both hardware setups' is not backed by data. A small table or correlation coefficient showing per-team ranks on each machine would support this claim.","section":"III"}],"recommendation":"major_revision","confidential_remarks":"The affiliation overlap between the top three teams and the author list is the most serious issue in this manuscript and should have been disclosed at submission. In addition to the revisions requested in the major comments, the editor may wish to ensure that the final version includes a clear conflict-of-interest statement, since the paper evaluates the authors' own competition entries without acknowledgment."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nI read the CARIC benchmark paper. The benchmark itself is a real and useful artifact: it defines a clear problem—heterogeneous LiDAR/camera teams inspecting structures without prior maps—implements it in Gazebo/RotorS with public code, and specifies evaluation metrics (blur, resolution, LoS) that are concrete enough to be reimplemented. Running the competition on two hardware setups with consistent rankings is a nice robustness check. That portion deserves credit.\n\nThe soft spot is the lessons-learned section. The top three teams are not external teams: KIOS CoE is the University of Cyprus group on the author list, XXH is NTU (also authors), and STAR is Sun Yat-sen University (again authors). The paper nowhere discloses that the analyzed teams are, in effect, the organizers and their close collaborators. That does not make the benchmark worthless, but it undermines the claim that 'no single approach has achieved robust performance across all scenarios' as a finding about the broader research community. Without per-run score data (only box plots are shown) or any baseline comparison, the reader cannot separate the organizers' performance from independent participants'. The qualitative lessons about exploration vs. inspection and task allocation are plausible and interesting, but they read as case studies of three systems built in-house, not as a neutral empirical evaluation.\n\nA secondary concern is that the scoring model, especially the motion-blur and resolution metrics, is not calibrated to real inspection outcomes. That is a minor issue for a simulation benchmark, but it limits how far the rankings can be interpreted.\n\nMy recommendation: the paper deserves peer review, but it needs major revision. At minimum, the authors must disclose the affiliation overlap and release per-run scores so the community can assess the empirical claims. If they re-frame the lessons as 'what we learned from our own teams,' that would be honest and still valuable.\n\nWould I bring it to a reading group? Maybe, mainly to discuss the conflict-of-interest issue in competition-based benchmarking. I'd cite the benchmark if I worked in multi-UAV inspection.","headline":"Useful benchmark artifact, but the lessons-learned section is compromised by undisclosed author-team overlap.","tokens_in":11004,"tokens_out":2024,"would_cite":true,"duration_ms":19348,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that CARIC, a simulation benchmark for heterogeneous multi-UAV inspection, provides a ready-to-use platform for developing and comparing planning algorithms, and that its first competition shows no single approach yet…","keywords":["multi-UAV inspection","heterogeneous robot team","motion planning benchmark","task allocation","line-of-sight communication","inspection image quality","aerial robotics","competition lessons"],"falsifier":"Run the same algorithms in the field on a structure with known defect locations, have experts label which captured images are usable for defect detection, and check whether the order of teams by CARIC score matches the order by expert-rated detection accuracy; a mismatch would refute the benchmark's validity.","tokens_in":10027,"feed_emoji":"🚁","tokens_out":11928,"duration_ms":180059,"temperature":0.7,"pith_summary":"The paper introduces CARIC, a simulation benchmark in which a heterogeneous team of LiDAR-equipped 'explorer' drones and camera-only 'photographer' drones must inspect unknown industrial structures from only bounding-box hints, under line-of-sight communication and a score based on image blur and resolution. The authors ran it as a competition at an international control conference and analyzed the top three entries. Their central claim is that CARIC is a ready-to-use, realistic testbed for multi-UAV task allocation and motion planning, and that the first competition demonstrates the problem's difficulty: no submitted approach performed robustly across all three scenarios. The value of this claim, if true, is that cooperative inspection without a prior structural model now has a shared platform for comparison and progress.","feed_headline":"No single strategy wins every scene in drone inspection benchmark","feed_subtitle":"A drone-inspection competition shows exploration speed, image quality, and task allocation each win different scenarios.","key_machinery":"The central object is the benchmark itself: a simulation stack built on a widely used open-source robotics simulator, with a heterogeneous fleet, a line-of-sight communication router that drops messages between occluded drones, and an evaluation function $Q = \\sum_{i} \\max_j q_{i,j}$, where $q_{i,j} = \\max_k (q_{\\mathrm{seen}} \\cdot q_{\\mathrm{blur}} \\cdot q_{\\mathrm{res}})$. The load-bearing identity is the multiplicative quality score: motion blur depends on pixel displacement during exposure (Equation 3), and spatial resolution is expressed as millimeter-per-pixel compared with a desired threshold (Equation 6). This score is what forces algorithms to balance speed, viewpoint distance, and camera aim, and the absence of any prior map is what makes exploration, mapping, and inspection need to be coordinated online under intermittent line-of-sight communication.","core_discovery":"CARIC operationalizes heterogeneous multi-UAV inspection as a scoring problem: explorers with LiDAR plus a camera build an online point-cloud map of structures enclosed by given bounding boxes, while photographers with only a camera fly to viewpoints to capture interest points; each image earns a quality score $q = q_{\\mathrm{seen}} \\cdot q_{\\mathrm{blur}} \\cdot q_{\\mathrm{res}}$, where $q_{\\mathrm{seen}}$ requires line-of-sight and field-of-view, $q_{\\mathrm{blur}}$ penalizes pixel motion during exposure (Eq. 3), and $q_{\\mathrm{res}}$ penalizes spatial resolution coarser than a desired millimeter-per-pixel (Eq. 6). The total score is the sum over interest points of the best image any drone captured. The authors' finding from the first competition is that the three top solutions embody a fundamental trade-off between exploration thoroughness and inspection time, between fast continuous scanning and sharp stop-and-scan imaging, and between simple volume-based task partitioning and workload-aware allocation, and that no single design wins all scenarios.","pith_inferences":["The multiplicative score could be used as a differentiable objective in trajectory optimization: since $q_{\\mathrm{blur}}$ depends on image-plane velocity, a planner could trade speed against expected blur directly rather than using waypoint stops as a proxy.","The benchmark's line-of-sight communication model invites communication-aware task allocation as an explicit scoring axis, such as counting messages or map freshness, which the current score ignores.","If the simulated blur and resolution metrics are not calibrated against real defect-detection performance, CARIC rankings may not transfer to field deployments; a validation study comparing CARIC scores with expert crack-detection rates on captured images would settle this.","The observation that volume-based workload estimation fails suggests inspection workload should be measured by surface area and viewpoint count, a metric the paper's own analysis supports but does not formalize as a benchmark statistic."],"forward_implications":["Treating exploration and inspection as separate sequential phases can leave photographers idle for long stretches; the fastest mapper in the competition won two scenarios but missed thin structures in a third.","Task allocation based on bounding-box volume is a poor workload proxy; the team that allocated by number of inspection viewpoints achieved more consistent coverage.","Continuous fast camera motion detects more candidate interest points, but the blur metric penalizes it heavily, so stop-and-scan yields lower recall but higher per-point quality and suggests role specialization among photographers.","Run-to-run variance of over 1000 score points in some scenarios shows that robustness of pathfinding and initialization is as important as the planning strategy itself.","No single submitted approach dominated all three scenes, so the benchmark can differentiate future multi-UAV inspection planners."],"supporting_citations":[{"why":"It supplies the modular quadrotor simulator that serves as the benchmark's physics and sensor ground truth.","marker":"[12]"},{"why":"It defines the LiDAR-equipped autonomous inspection UAV whose design motivates the explorer role in the fleet.","marker":"[13]"},{"why":"It provides the in-flight image quality check concept from which the motion blur metric is adapted.","marker":"[14]"},{"why":"It provides the ground sampling distance formulation used for the spatial resolution metric.","marker":"[15]"},{"why":"It inspires the distributed mTSP-based path-planning method used by the first-place approach.","marker":"[16]"},{"why":"It is the frontier-based exploration framework adopted by one of the teams for explorer mapping.","marker":"[17]"},{"why":"It is the hierarchical viewpoint clustering and planning method used by one team for photographer inspection.","marker":"[18]"},{"why":"It supplies the incremental map-sharing mechanism used by one team's exploration strategy.","marker":"[19]"},{"why":"It provides the smooth collision-free trajectory generation used for local planning.","marker":"[20]"}],"fun_headline_variants":["Drone benchmark shows no universal inspection strategy","Heterogeneous drone teams: trade-offs, not one best plan","CARIC competition reveals exploration vs. imaging trade-off","No one drone strategy tops all inspection scenarios","Inspection drones: speed, sharpness, allocation compete"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's ranking assumes the simulated line-of-sight, motion-blur, and spatial-resolution metrics faithfully capture whether images are good enough for real structural defect analysis; until those metrics are calibrated against real inspection outcomes, a high CARIC score may not mean a high real-world inspection quality.","fun_headline_variants_meta":{"raw":{"variants":["Drone benchmark shows no universal inspection strategy","Heterogeneous drone teams: trade-offs, not one best plan","CARIC competition reveals exploration vs. imaging trade-off","No one drone strategy tops all inspection scenarios","Inspection drones: speed, sharpness, allocation compete"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000143,"raw_usage":{"total_tokens":1150,"prompt_tokens":904,"completion_tokens":246,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":170}},"tokens_in":520,"tokens_out":246,"duration_ms":2938,"temperature":1.0,"reasoning_tokens":170,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:57:19.942443+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same algorithms in the field on a structure with known defect locations, have experts label which captured images are usable for defect detection, and check whether the order of teams by CARIC score matches the order by expert-rated detection accuracy; a mismatch would refute the benchmark's validity.","supporting_citations":[{"cited_title":"Rotors—a modular gazebo mav simulator framework,","cited_arxiv_id":null,"evidence_quote":"It supplies the modular quadrotor simulator that serves as the benchmark's physics and sensor ground truth."},{"cited_title":"Autonomous underground flight with m300 rtk and the emesent hovermap","cited_arxiv_id":null,"evidence_quote":"It defines the LiDAR-equipped autonomous inspection UAV whose design motivates the explorer role in the fleet."},{"cited_title":"Rapid in-flight image quality check for uav-enabled bridge inspection,","cited_arxiv_id":null,"evidence_quote":"It provides the in-flight image quality check concept from which the motion blur metric is adapted."},{"cited_title":"Evaluation and enhancement of resolution-aware coverage path planning method for surface inspection using unmanned aerial vehicles,","cited_arxiv_id":null,"evidence_quote":"It provides the ground sampling distance formulation used for the spatial resolution metric."},{"cited_title":"Swarm path planning for the deployment of drones in emergency response missions,","cited_arxiv_id":null,"evidence_quote":"It inspires the distributed mTSP-based path-planning method used by the first-place approach."},{"cited_title":"Fuel: Fast uav exploration using incremental frontier structure and hierarchical planning,","cited_arxiv_id":null,"evidence_quote":"It is the frontier-based exploration framework adopted by one of the teams for explorer mapping."},{"cited_title":"Star-searcher: A complete and efficient aerial system for autonomous target search in complex unknown environments,","cited_arxiv_id":null,"evidence_quote":"It is the hierarchical viewpoint clustering and planning method used by one team for photographer inspection."},{"cited_title":"Racer: Rapid collaborative exploration with a decentralized multi-uav system,","cited_arxiv_id":null,"evidence_quote":"It supplies the incremental map-sharing mechanism used by one team's exploration strategy."}],"review_version":1}