{"id":"4ecd7661-944a-45a0-b480-5434218936b5","arxiv_id":"2506.04438","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 2019 survey of 110 experts shows autonomous systems practice relied mostly on rule-based methods, high autonomy levels, and relatively little machine learning, with weak adoption of modeling standards and certification.","lead":"This paper reports results from a 2019 survey of 110 software developers and researchers about how autonomous systems are actually built, covering domains, tools, standards, and testing practices. It offers one of the few data points on industrial practice in a field dominated by literature reviews and proprietary case studies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The state-of-the-practice claim rests on a convenience sample of 110 respondents, with 69% from space/aviation/military domains and only 9% automotive; the paper's own Section 5 disclaimer contradicts the C1 generalization, so the headline percentages may reflect the recruitment network rather…","rationale":"The reader's weakest assumption, representativeness of the convenience sample, is exactly the load-bearing assumption for C1 and the headline percentages. The concern lands because the paper's own Section 5 admits the limitation while the abstract and Section 6 still assert a general state-of-the-practice. I agree with the CONDITIONAL verdict: the survey is a transparent, carefully described descriptive study of its 110 respondents, and the authors deserve credit for piloting the instrument and reporting per-question denominators. However, the specific percentages are sample statistics with wide error bars; the 'full autonomy' and 'ML usage' findings are based on 50 and 58 AUCs clustered in 38 respondents, and the domain mix is heavily space/aviation/military. A conditional acceptance requiring publication of the instrument and raw data, uncertainty intervals, and qualified generalizations directly addresses the issue. No internal inconsistency or computational error is apparent; the concern is external validity, not soundness. The concrete test of stratified re-analysis would settle whether the domain skew actually changes the conclusions.","tokens_in":21629,"tokens_out":6942,"duration_ms":62824,"concrete_test":"Ask the authors to release the de-identified raw survey responses (including the domain field and the mapping from respondents to AUCs) and re-run the headline analyses (RQ2b full-autonomy percentages from Figure 7, RQ2c ML usage of 12%, and RQ4c certification rate) stratified by domain (space/aviation/military vs. all other domains) and with cluster-robust 95% confidence intervals at the respondent level for all AUC-level estimates. If the stratified percentages differ from the pooled percentages by more than 10 percentage points, or if any confidence interval half-width exceeds +/-15 percentage points, the pooled state-of-the-practice percentages are not robust to the sample's domain skew and the C1 generalization fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution C1 ('established the state-of-the-practice of developing autonomous systems') and the Section 6 conclusion ('safety-critical systems with a substantial degree of autonomy have been successfully developed in the industry') generalize from 110 self-selected respondents to the entire industry. This generalization is not supported. The sample was recruited by convenience sampling and snowballing, including 'over 300 authors of research papers related to autonomy,' and the author team is NASA-affiliated; 40% of respondents came from space, 16% from aviation, and 13% from military, leaving only 9% from automotive, the sector with the largest commercial investment in autonomy. The authors' own earlier MBSwE survey [24] reported 33% automotive, showing how recruitment channels change domain mix. Key percentages rest on very small and clustered denominators: the full-autonomy figures (50-67%) are computed over 50 AUCs supplied by 38 respondents, with multiple AUCs per respondent and no cluster adjustment; the ML-usage figure of 12% is 7 out of 58 AUCs; the security-V&V figure of 11% is 4 out of 36 respondents. Section 5 explicitly disclaims that 'the results based on one survey would be valid for all autonomous systems,' directly contradicting the state-of-the-practice framing in C1 and the abstract. If the sample skews toward NASA-affiliated space/defense projects, the headline percentages (full autonomy 50-67%, ML 12%, certification 24%) may characterize that subpopulation, not the industry-wide state of practice.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the first wave of a longitudinal survey-based study of autonomous-systems development, based on an anonymous online survey administered in 2019. Of 129 respondents who started the survey, 110 reported having worked on autonomous systems and/or MBSwE and answered questions in six blocks covering industry domains, safety criticality, programming languages, code origins, autonomy details for up to 58 autonomous components (AUCs), reuse, processes and standards, verification and validation, and bugs. The main reported findings are a space/aviation/military-dominated respondent pool, high safety criticality, C/C++ as dominant languages, rule-based rather than ML algorithms (ML used for 12% of AUCs), full computer autonomy for 50–67% of AUCs across four autonomy tasks, low incidence of certification, and sparse V&V of security and reused artifacts. The paper claims four contributions: establishing the state-of-the-practice (C1), quantifying autonomy and reuse benefits and challenges (C2), identifying processes and standards (C3), and exploring V&V (C4), followed by practitioner recommendations.","tokens_in":21936,"tokens_out":5385,"duration_ms":46479,"significance":"If the descriptive results are accepted as an accurate snapshot of the 2019 respondent population, the paper is a valuable empirical complement to the many literature-review studies in this area. Its strengths are the structured survey design following Kitchenham and Pfleeger, the explicit reporting of the number of respondents behind each figure and table, the pilot study, the anonymous administration, and the stated intent to repeat the survey longitudinally. The usefulness of the findings as evidence about industry practice in general, however, is limited by the convenience sample, the domain skew toward space, aviation, and military, and the small denominators behind several headline percentages; these factors must be addressed before the 'state-of-the-practice' claim can be considered established.","major_comments":[{"comment":"The central claim that the survey 'established the state-of-the-practice of developing autonomous systems' (C1) and the Section 6 conclusion that 'safety-critical systems with a substantial degree of autonomy have been successfully developed in the industry' generalize beyond the respondent pool, but the paper explicitly disclaims such generalization in Section 5: 'we cannot claim that the results based on one survey would be valid for all autonomous systems.' The sample was recruited via non-probabilistic convenience sampling and snowballing (Section 3), including contacts of the NASA-affiliated author team and over 300 authors of autonomy-related papers, and the respondent mix is heavily skewed toward space (40%), aviation (16%), and military (13%), with only 9% automotive (Figure 2). The authors' own prior MBSwE survey [24] had 33% automotive, which the paper itself cites as evidence that recruitment channels change the domain mix. Please reframe the contributions and conclusions as describing the surveyed sample, or provide substantive evidence of representativeness (e.g., comparison with known population demographics, weighting, or sensitivity analysis).","section":"Section 1 (C1), Section 3, Section 5, Section 6"},{"comment":"Key percentages are computed over very small and non-independent units: data about 58 AUCs were provided by 38 respondents, with 28 respondents reporting one AUC and the rest reporting multiple AUCs. Figure 7 is based on 50 AUCs and Figure 8 on 58 AUCs, so the headline full-autonomy figures (50–67%) and the ML-usage figure of 12% (7 of 58 AUCs) have wide sampling uncertainty and no adjustment for clustering of AUCs within respondents. Section 5's statement that 'these are fairly large sample sizes that allow drawing valid conclusions' is difficult to support for denominators of 29–58. Please report raw counts and confidence intervals for all key proportions, and either account for the clustering or qualify the precision of the estimates.","section":"Section 4.2, Figures 7–8, Section 5"},{"comment":"The paper reports two different numbers for V&V of security: Section 4.5.1 states 'only 11% of respondents verified and validated security' (about 4 of the 36 respondents who answered that question), while the final recommendations state 'Verification and validation of security was rare, done by only 4% of the respondents.' These cannot both be correct. Please reconcile the discrepancy and ensure Table 9 matches the detailed results.","section":"Section 4.5.1 vs. Section 6 recommendations"}],"minor_comments":[{"comment":"The text contains a typo: 'concpets' should be 'concepts'.","section":"Section 1"},{"comment":"The description of Taylor et al. uses 'reproducability' and should be 'reproducibility'.","section":"Section 2, reference [32]"},{"comment":"The label 'Austonomous response and adaptation of mission' should read 'Autonomous response and adaptation of mission'.","section":"Figure 23"},{"comment":"The sentence 'Humans are know to have evaluation apprehension' should read 'Humans are known to have evaluation apprehension'.","section":"Section 5"},{"comment":"Several bar charts report only percentages without raw counts; adding the underlying counts would improve interpretability, especially given the small denominators discussed in the major comments.","section":"Figures 7–16"}],"recommendation":"major_revision","confidential_remarks":"The paper's honest threats-to-validity section is a strength, but the abstract and contribution statements outrun the caution expressed there. I would ask the authors to align the framing of the 'state-of-the-practice' claim with the actual sampling method before publication; this is fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a genuinely useful descriptive dataset, and the paper reports it transparently. It is the only multi-domain survey I know that combines autonomy, MBSwE, and reuse in one instrument, and the 2019 snapshot is clearly labeled as such. The headline numbers—12% ML usage, 50-67% full autonomy across tasks, 25% reuse-specific bugs, only 24% certification—are real observations about the respondents and will be cited, I suspect, as a baseline for what autonomy practice looked like before the current ML wave. The authors are careful to annotate every figure with the actual number of respondents, and the Section 5 threats-to-validity discussion is honest, not boilerplate.\n\nThe soft spots are real but not disqualifying. The paper repeatedly claims to have 'established the state-of-the-practice of developing autonomous systems,' both in C1 and in the Section 6 conclusion about safety-critical systems being 'successfully developed in the industry.' That generalization goes beyond what the data support. The sample is a convenience/snowball sample of 110 respondents, 40% from space, 16% aviation, 13% military, only 9% automotive, recruited in part by emailing over 300 authors of autonomy papers. The NASA affiliation of the author team plausibly shaped the network. The authors even admit in Section 5 that they 'cannot claim that the results based on one survey would be valid for all autonomous systems,' which directly contradicts the C1/abstract framing. Also, several percentages rest on very small denominators: the full-autonomy figures come from 50 AUCs supplied by 38 respondents, the ML figure is 7 out of 58 AUCs, and the security-V&V figure is 4 out of 36 respondents. No confidence intervals, no cluster adjustment for multiple AUCs per respondent. For a descriptive survey this is acceptable if the claims are framed as sample-specific; it is not acceptable if the claim is industry-wide state of practice.\n\nWhat would make this publishable without overreach: publish the survey instrument and cleaned response data, add a brief uncertainty analysis (or at least consistently use 'in our sample' language), and demote the 'state-of-the-practice' claim to 'state of our sample.' The paper deserves a serious referee; the data are new and the topic matters. I'd send it to review, but the reviewers should push on the generalization.","headline":"A transparent, genuinely useful 2019 snapshot of autonomy development practice, but the 'state-of-the-practice' framing overreaches its convenience sample.","tokens_in":22470,"tokens_out":2180,"would_cite":true,"duration_ms":20889,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 2019 survey of 110 experts finds that autonomous systems in practice ran mostly on rule-based code, not machine learning.","keywords":["autonomous systems","survey study","state of the practice","verification and validation","model-based software engineering","software reuse","machine learning","autonomous components"],"falsifier":"Repeating the survey with a sample that is randomly drawn or balanced across industry domains and finding markedly different rates of full autonomy or machine learning usage would falsify the state-of-the-practice claim. A simpler check is whether the 40% space share of respondents matches the actual distribution of autonomous-systems development effort across sectors.","tokens_in":21456,"feed_emoji":"🤖","tokens_out":4575,"duration_ms":38646,"temperature":0.7,"pith_summary":"This paper tries to establish the state of the practice in developing autonomous systems as of 2019, using an anonymous online survey answered by 110 experts from industry and academia. It claims that high levels of autonomy were already common: for each of four autonomy tasks, 50% to 67% of autonomous components operated with full computer autonomy. It also claims that development was still dominated by traditional approaches, with rule-based algorithms used for 35% of components and machine learning for only 12%. The authors present these results as the first part of a longitudinal study meant to track how autonomous systems development evolves.","feed_headline":"2019 survey: autonomy was common, machine learning rare","feed_subtitle":"110 experts reported full computer autonomy in 50–67% of autonomous components, but only 12% used machine learning.","key_machinery":"The central object is the Autonomous Component (AUC), defined as a component that provides autonomous capabilities and supports autonomous operations; respondents provided details for 58 AUCs. The argument is carried by a 48-question branching online survey, administered from April 25 to June 20, 2019, whose sections cover domain, safety criticality, programming languages, autonomy levels, reuse, processes and standards, verification and validation, and bugs. The survey's four-level autonomy scale and its four information-processing tasks, adapted from prior frameworks, produce the headline percentages.","core_discovery":"The paper's central claim is that in 2019, industrial practice in autonomous systems was more conventional than media attention to machine learning suggested. Across the four information-processing tasks used in the survey, between 50% and 67% of autonomous components had full computer autonomy, and the most common development methods were rule-based algorithms, planning systems, and statistical filtering. Only 9% of components used offline machine learning and 3% used online machine learning. The authors conclude that safety-critical systems with substantial autonomy were being built successfully with existing software technology, while naming system complexity, environmental uncertainty, and achieving the desired level of autonomy as the principal challenges.","pith_inferences":["Because 40% of respondents came from the space domain and another 29% from aviation or military, the 'state of the practice' likely describes safety-critical aerospace practice more accurately than it describes consumer automotive or service robotics; this is an inference from the reported domain distribution.","The sharpest test of the paper's claims will come from the planned second survey: if machine learning usage and certification rates have not moved by the mid-2020s, the 2019 baseline will look less like a lag and more like a persistent feature of the field.","The reuse findings hint that conventional assumptions about reuse lowering defect rates may not transfer to autonomous systems, because reuse-specific bugs were nearly as common as autonomy-specific bugs; the paper does not state this directly."],"forward_implications":["If these numbers describe 2019 practice, then full autonomy was not a distant prospect; it was already deployed across a majority of surveyed autonomous components.","If machine learning was used in only 12% of components, then assurance methods aimed at learning-enabled systems were addressing a small slice of the autonomous systems actually being built at that time.","If only 24% of systems went through certification and 26% of autonomous components were part of a certified system, then certification was a clear gap in 2019 practice.","If 25% of respondents reported reuse-specific bugs while only 19% verified reused artifacts, then reuse without dedicated verification and validation is a concrete risk in autonomous systems development."],"supporting_citations":[{"why":"Supplies the four information-processing task categories and the human-computer level-of-autonomy scale used to frame the survey's autonomy questions.","marker":"[7]"},{"why":"Provides the earlier model-based software engineering survey whose respondent base and findings this survey extends and contrasts.","marker":"[24]"},{"why":"Gives the earlier space-domain survey result that lower autonomy levels dominated, against which this paper's higher autonomy levels are compared.","marker":"[25]"},{"why":"Provides the stepwise survey design and administration methodology the authors followed.","marker":"[38]"},{"why":"Supplies the authors' prior bug study used to contextualize autonomy-specific bugs and the paper's recommendations.","marker":"[34]"},{"why":"Defines the safety criticality levels and the reusable software component concept that the survey uses.","marker":"[9]"}],"fun_headline_variants":["2019 autonomy: rule-based methods dominate, ML rare","Survey: 2019 autonomous development shuns machine learning","Industrial autonomy in 2019: conventional, not ML","2019 industrial autonomous systems: ML use minimal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 110 respondents, recruited through convenience sampling and snowballing, are representative of the broader population of autonomous systems developers, despite the sample's heavy skew toward space, aviation, and military domains.","fun_headline_variants_meta":{"raw":{"variants":["2019 autonomy: rule-based methods dominate, ML rare","Survey: 2019 autonomous development shuns machine learning","Industrial autonomy in 2019: conventional, not ML","2019 industrial autonomous systems: ML use minimal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1334,"prompt_tokens":841,"completion_tokens":493,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":428}},"tokens_in":457,"tokens_out":493,"duration_ms":7097,"temperature":1.0,"reasoning_tokens":428,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:42:02.450797+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeating the survey with a sample that is randomly drawn or balanced across industry domains and finding markedly different rates of full autonomy or machine learning usage would falsify the state-of-the-practice claim. A simpler check is whether the 40% space share of respondents matches the actual distribution of autonomous-systems development effort across sectors.","supporting_citations":[{"cited_title":"A model for types and levels of human interaction with automation,","cited_arxiv_id":null,"evidence_quote":"Supplies the four information-processing task categories and the human-computer level-of-autonomy scale used to frame the survey's autonomy questions."},{"cited_title":"Sur- vey on Model-Based Software Engineering and Auto-Generated Code,","cited_arxiv_id":null,"evidence_quote":"Provides the earlier model-based software engineering survey whose respondent base and findings this survey extends and contrasts."},{"cited_title":"Intelligent control for spacecraft autonomy - an industry survey,","cited_arxiv_id":null,"evidence_quote":"Gives the earlier space-domain survey result that lower autonomy levels dominated, against which this paper's higher autonomy levels are compared."},{"cited_title":"Personal opinion surveys,","cited_arxiv_id":null,"evidence_quote":"Provides the stepwise survey design and administration methodology the authors followed."},{"cited_title":"The anatomy of software changes and bugs in autonomous operating system,","cited_arxiv_id":null,"evidence_quote":"Supplies the authors' prior bug study used to contextualize autonomy-specific bugs and the paper's recommendations."},{"cited_title":"DO-178C/ED-12C: Software considerations in airborne systems and equipment certification,","cited_arxiv_id":null,"evidence_quote":"Defines the safety criticality levels and the reusable software component concept that the survey uses."}],"review_version":1}