{"id":"2e5c1af1-b5e1-482c-944a-095a7e4d370c","arxiv_id":"2505.03472","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of 97 CAV simulators and testbeds that derives eight software requirements and four testbed-selection recommendations for moving from simulation to reality.","lead":"This paper surveys software frameworks and testbeds used to develop connected and automated vehicles, from simulators to full-scale cars. It maps which technologies fit each testing stage and lists the key software requirements for moving between stages.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey conclusions rest on an undocumented and apparently skewed testbed sample, and Section 4.1 contains an internal count inconsistency that undermines the inductive basis for the findings.","rationale":"The paper is a competent taxonomy of simulators, small-scale testbeds, and full-scale testbeds, and its requirement list is useful as a checklist. However, its central claims are empirical generalizations: the conclusion that testbed choice is 'largely driven by simplicity and convenience' and the four findings in Section 8.2 depend on the representativeness of the inventory. The weakest link is the unstated sampling frame. The internal count discrepancy (45 ROS / 4 custom / 20 no-middleware versus Table 2's 54 entries with 40 / 4 / 10) is a concrete symptom that the data were assembled without a formal protocol. The reader's CONDITIONAL verdict already captures this concern; I agree with it and see no reason to change it. The condition should be enforced: the authors should document the selection process and disclose overlap with their own prior survey and testbed.","tokens_in":39334,"tokens_out":5049,"duration_ms":51503,"concrete_test":"Re-run the survey with a documented protocol: specify databases (e.g., IEEE Xplore, Scopus), a time window (e.g., 2015-2025), and explicit inclusion/exclusion rules; then recompute Table 2 and Finding 2 twice, once on the full result set and once after removing entries with any author overlap with this paper. If the recommendation that planning and control are best tested on small-scale platforms weakens or flips in the non-overlapping subset, the current conclusions are artifacts of sample selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 9's conclusion and Findings 2 and 3 generalize from an inventory whose selection procedure is never specified: Sections 3-5 state only that simulators and testbeds were 'found' during literature review, with no search window, databases, or inclusion criteria. The sample is also internally inconsistent: Section 4.1 reports 45 ROS-based small-scale testbeds, 4 custom-middleware testbeds, and 20 testbeds without middleware, but Table 2 contains 54 testbeds and, by the table's own middleware column, only 40 ROS, 4 custom, and 10 without middleware. This is not a minor typo, because the paper's core claims ('ROS dominates small-scale testbeds'; 'planning and control are best tested on small-scale platforms') are inductions from this table. The inventory also overlaps heavily with the authors' own CPM Lab and prior survey [166], so the apparent prevalence of ROS-based academic testbeds may reflect the authors' community rather than the field. If the sample is biased, the 'choice driven by simplicity and convenience' conclusion and the four findings in Section 8.2 do not have a valid inductive base.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews software frameworks, middlewares, simulators, and testbeds for connected and automated vehicles (CAVs). It proposes a seven-way simulator taxonomy, catalogs 30 simulators, 54 small-scale testbeds, and 13 full-scale testbeds, and derives eight requirements for software under test when moving from simulation to small-scale to full-scale. The paper concludes that the choice of architecture and testbed scale is largely driven by simplicity and convenience, and it makes four recommendations, e.g., planning and control should be tested on small-scale testbeds while traffic management should be evaluated in simulation.","tokens_in":39564,"tokens_out":5597,"duration_ms":47949,"significance":"If the inventory were representative, the paper would provide a useful map for selecting testbed scales and middlewares, and the requirements checklist could inform the design of new CAV software frameworks. The survey's strengths are its breadth and organization, and the distinction between API-based and middleware-based simulation architectures is helpful. However, the central findings are inductions from an inventory whose construction is not documented and whose internal counts are inconsistent; these issues must be resolved before the conclusions can be accepted as general.","major_comments":[{"comment":"The text reports that 45 of the considered small-scale testbeds use ROS, 4 use custom middleware, and 20 do not use middleware, but Table 2 contains 54 testbeds and its middleware column lists 40 ROS-based, 4 custom, and 10 without middleware. The numbers are inconsistent in both readings (45+4+20=69 vs. 54; 40+4+10=54), and because the 'ROS dominates' claim and Finding 2 rely on this distribution, the counts must be reconciled and the affected claims re-derived.","section":"Section 4.1 and Table 2"},{"comment":"No search protocol, inclusion/exclusion criteria, database list, or time window is given for the inventory of simulators and testbeds. Since Findings 2-4 and the Section 9 conclusion generalize from this sample, the inductive basis is not verifiable; please add a methodology subsection and discuss potential selection bias, including the over-representation of the authors' own CPM Lab and related work in Table 2 and Section 8.2.","section":"Sections 3-5"},{"comment":"The recommendations, e.g., 'planning and control are best tested on a small-scale testbed,' are asserted as conclusions of the survey, but the link from the tabulated data to these preferences is never made explicit. In particular, Table 2 contains many small-scale testbeds for planning and control, yet the survey does not compare success metrics or costs against full-scale alternatives; please state which evidence in Tables 1-3 supports each finding, or present the findings as informed opinions rather than data-driven conclusions.","section":"Section 8.2 and Section 9"}],"minor_comments":[{"comment":"The caption of Fig. 3 states 'We found four classifications,' while the text and Table 1 present seven categories; the caption should be corrected.","section":"Section 3.2 and Fig. 3"},{"comment":"The introductory paragraph says 'we summarize these recommendations in five findings in Section 8.2,' but Section 8.2 lists four findings; please align the counts.","section":"Section 8"},{"comment":"The table caption reads 'Overview of Overview of the full-scale testbeds'; remove the duplicated phrase.","section":"Section 5.1 and Table 3"},{"comment":"The text says testbeds are grouped into 9 categories, Fig. 4's caption lists 6 or 7 categories, and Table 2 shows 7 labeled groups plus education/communication entries; the category count should be made consistent across text, figure, and table.","section":"Section 4.1 and Fig. 4"},{"comment":"The Year column should state whether it records first release, cited version, or last update; for example, LGSVL is listed as 2020 despite earlier releases, and the 'None' middleware entries could be clarified as 'not specified' rather than 'no middleware'.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The authors should be aware that heavy self-citation in the inventory tables and recommendations may raise independence concerns among reviewers; using an external validation of the testbed selection or explicitly discussing the authors' relationship to CPM Lab would strengthen the manuscript. The paper fits the journal's scope, but the missing methodology and the count inconsistency are the primary barriers to acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a genuinely useful survey but the numbers don't hold up. The paper claims 45 ROS-based small-scale testbeds, 4 custom, and 20 without middleware, but Table 2 — the paper's own table — shows 41 ROS, 3 custom, and 10 without. That is not a minor typo, because the findings in Section 8.2 are inductions from this table.\n\nWhat's good: the survey is well-organized and covers a lot of ground. Thirty simulators, 54 small-scale testbeds, 13 full-scale testbeds, with middleware, focus, and publication data. The eight requirements (real-time, determinism, distributed systems, etc.) are a reasonable checklist, and the discussion of how AUTOSAR Adaptive, ROS 2, and DDS handle them is genuinely informative. Adding E/E and component simulator categories to Li et al.'s taxonomy is a small but sensible extension.\n\nThe soft spots are real. The selection method is never described: no search window, databases, or inclusion criteria. The tables are just 'what we found during our literature review.' That alone makes the prevalence claims fragile. On top of that, the count mismatch shows the table wasn't carefully checked. And the paper leans heavily on the authors' own CPM Lab and prior survey [166] in both the tables and the recommendations. Self-citation isn't inherently bad, but when the sample is undocumented, the overlap becomes a legitimate concern.\n\nI should say the central argument — that requirements get stricter as you move from simulation to physical testbeds — holds up. It's a sensible qualitative claim, and the requirement sections support it. The problem is the paper then goes further and makes quantitative claims about what the field does ('ROS dominates', 'most testbeds are X') without a defensible sample.\n\nSo: not a new research result, and the inductive base needs work. But for someone entering CAV testing and wanting a map of simulators and testbeds, this would be a helpful starting point after revision. I'd send it to review with a demand for a proper methodology section and corrected counts. Right now I would not rely on its prevalence statistics, but the taxonomy and requirement analysis have value.","headline":"Useful survey of CAV testbeds and middlewares, but the counts in Section 4.1 don't match Table 2, so the inductive findings need a documented sample and correction before they can be trusted.","tokens_in":40055,"tokens_out":4581,"would_cite":false,"duration_ms":39136,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Moving automated-vehicle software from simulation to full road tests is a ladder of eight requirements, and most testbed choices track convenience.","keywords":["connected and automated vehicles","software architecture","automotive middleware","simulation","small-scale testbeds","full-scale testbeds","real-time systems","testbed selection"],"falsifier":"A reproducible re-survey with explicit inclusion criteria that adds industrial and non-European testbeds would settle the inventory's representativeness; if it shows testbed and middleware choices tracking funding or regulation rather than simplicity, the paper's main conclusion is wrong. A second check is to look for a functioning full-scale CAV testbed whose software stack intentionally lacks one of the eight stated requirements and still passes safety review.","tokens_in":39158,"feed_emoji":"🚗","tokens_out":8070,"duration_ms":72022,"temperature":0.7,"pith_summary":"This survey argues that connected and automated vehicle (CAV) software is validated in three stages—pure simulation, small-scale testbeds with miniature vehicles, and full-scale testbeds with real cars—and that each transition imposes a new set of software requirements. It catalogs 30 simulators, 54 small-scale testbeds, and 13 full-scale testbeds, and explains the middleware choices that appear at each scale. The paper's central conclusion is that the choice of testbed scale and software architecture is largely driven by simplicity and convenience rather than by a rigorous method. On that basis it issues four concrete recommendations, including testing planning and control at small scale and testing traffic management in simulation. A reader comes away with a checklist of eight requirements any CAV software framework must satisfy before moving toward road-ready vehicles.","feed_headline":"Survey: 8 requirements gate the road from simulation to self-driving","feed_subtitle":"It maps simulators, scale-model fleets, and full-scale vehicles, then finds most choices track convenience.","key_machinery":"The load-bearing structure is a requirements ladder: a set of eight software requirements ordered by testbed scale, starting with energy efficiency and real-time execution at the small scale and ending with orchestration, safety and security, platform compatibility, and experiment recording at full scale. The ladder is built by comparing middleware use across 30 simulators, 54 small-scale testbeds, and 13 full-scale testbeds, grouped into seven simulator classes, nine small-scale categories, and three full-scale architecture types. It is the mechanism that connects observed testbed practice to the paper's recommendations.","core_discovery":"The paper's core claim is that there is a systematic ladder from simulation to small-scale to full-scale testing, and the ladder is defined by progressively stricter software requirements. In simulation, execution can be paused or accelerated, failures are cheap, and computing resources appear abundant; small-scale physical platforms add the need for real-time execution, deterministic and repeatable behavior, energy awareness, and support for multiple distributed agents linked by wireless communication; full-scale vehicles add orchestration and resource management, safety and security, compatibility with heterogeneous automotive hardware and protocols, and complete experiment recording. The paper derives eight requirements in total and uses them to explain four findings: real-time behavior should be verified with dedicated real-time hardware and hardware-in-the-loop; planning, control, racing, and platooning are best evaluated on small-scale testbeds; traffic management should be tested in simulation; and edge computing and V2X communication are best tested at small or full scale. Underlying all of this is the observation that the choice of both architecture and testbed scale is largely driven by simplicity and convenience.","pith_inferences":["If testbed choice is indeed convenience-driven, a scoring tool that weighs the eight requirements against available testbed capabilities could replace ad hoc selection and is a direct next step the paper leaves implicit.","The eight requirements double as a maturity rubric: a software stack that cannot satisfy the full-scale requirements should not be presented as deployment-ready.","The survey's own inventory suggests that mixed-fidelity validation—simulation for traffic, small-scale for planning, full-scale for V2X—is the practical default, even though no single platform covers all needs.","A testable prediction follows from the convenience claim: testbeds whose platforms are harder to set up, for example those requiring custom middleware or strict real-time hardware, should be underused relative to their scientific value; a bibliometric analysis of which testbeds appear in method evaluations could check this."],"forward_implications":["Researchers evaluating planning, control, racing, or platooning algorithms should choose a small-scale testbed, moving to simulation only for traffic-level questions.","Claims about real-time behavior should be demonstrated on dedicated real-time hardware or with hardware-in-the-loop testing before small- or full-scale validation.","Traffic management and city-scale flow questions belong in simulation rather than on physical testbeds.","V2X and edge-computing experiments should begin with network simulators and small-scale proofs of concept before moving to full-scale on-road systems.","New CAV middleware frameworks should be designed against all eight requirements from the start, since later stages of testing will demand them."],"supporting_citations":[{"why":"Defines connected and automated vehicles, fixing the scope of the survey's claims.","marker":"[98]"},{"why":"Supplies the five simulator classes that the paper extends to seven.","marker":"[144]"},{"why":"Provides the small-scale testbed survey and the cost-safety rationale for scaled testing.","marker":"[166]"},{"why":"Defines the realism-cost trade-off and the testing taxonomy the paper builds on.","marker":"[217]"},{"why":"Compares major architecture platforms, underpinning the middleware analysis.","marker":"[105]"},{"why":"Introduces modern automotive middlewares, the core software concept examined.","marker":"[129]"},{"why":"Describes the design and real-time properties of a widely used middleware, referenced throughout the testbed tables.","marker":"[153]"},{"why":"Characterizes distributed automotive E/E architectures, grounding the platform-compatibility requirement.","marker":"[263]"},{"why":"Proposes an experimental architecture for CAV testbeds, used as evidence for deterministic behavior and recording needs.","marker":"[126]"},{"why":"Documents full-scale autonomous driving practices and risks behind the safety and cost requirements.","marker":"[254]"}],"fun_headline_variants":["8 requirements link CAV simulation to full-scale testing","From simulation to road: 8 software requirements for CAVs","CAV testbeds: convenience beats rigor, survey says","Eight requirements steer CAV testing from sim to scale","Sim-to-real CAV ladder: 8 requirements, convenience-driven choices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey assumes that its inventory of 30 simulators, 54 small-scale testbeds, and 13 full-scale testbeds is representative, even though it reports no search protocol, inclusion criteria, or time window, and the sample skews toward the authors' own testbeds and community.","fun_headline_variants_meta":{"raw":{"variants":["8 requirements link CAV simulation to full-scale testing","From simulation to road: 8 software requirements for CAVs","CAV testbeds: convenience beats rigor, survey says","Eight requirements steer CAV testing from sim to scale","Sim-to-real CAV ladder: 8 requirements, convenience-driven choices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000939,"raw_usage":{"total_tokens":3993,"prompt_tokens":904,"completion_tokens":3089,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":3005}},"tokens_in":520,"tokens_out":3089,"duration_ms":21597,"temperature":1.0,"reasoning_tokens":3005,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:50:21.080179+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reproducible re-survey with explicit inclusion criteria that adds industrial and non-European testbeds would settle the inventory's representativeness; if it shows testbed and middleware choices tracking funding or regulation rather than simplicity, the paper's main conclusion is wrong. A second check is to look for a functioning full-scale CAV testbed whose software stack intentionally lacks one of the eight stated requirements and still passes safety review.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the realism-cost trade-off and the testing taxonomy the paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Characterizes distributed automotive E/E architectures, grounding the platform-compatibility requirement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents full-scale autonomous driving practices and risks behind the safety and cost requirements."}],"review_version":1}