{"id":"04c3730b-ae85-4e67-82b8-8eaee999f933","arxiv_id":"2412.10476","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of 314 web application testing papers from 2014 to 2023, organized by test generation, execution, evaluation, tools, and open challenges.","lead":"This paper is a survey of 314 research papers on testing web applications, covering how test cases are generated, executed, and evaluated. It is useful as a reference map for researchers and practitioners who want a broad picture of the field's last decade.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's corpus-selection methodology (title-restricted queries, open-access filter, five publishers) is not validated against the field, so the reported trends may not be representative of WAT research 2014–2023.","rationale":"The reader's weakest assumption is the representativeness of the 314-paper corpus, and I agree this is the most load-bearing point. The survey's empirical content—every distribution, proportion, and trend—is a direct function of which papers were included. If the selection procedure biases the sample, the conclusions are unreliable no matter how carefully the included papers are summarized. I considered alternative concerns: (i) the scope contradiction where Section 1 lists regression testing and failure diagnosis as covered while Section 3.2 excludes them, and (ii) the denominator inconsistencies in Figures 5, 8, 9, and 10 where counts do not sum to 314. Both are real and should be fixed, but they are secondary: the scope issue can be addressed by explicitly narrowing the claimed scope, and the denominator issues are errors that can be corrected by re-labeling the base. The sample-representativeness issue cannot be patched cosmetically; it requires a validation study or a re-run of the search. The paper's own confidence statement in Section 3.2 is an assertion, not evidence. Hence the central claim remains conditional until a gold-standard comparison is performed. This is exactly the reader's point, so agreement is 'agree' and the verdict remains CONDITIONAL (no change).","tokens_in":47801,"tokens_out":8690,"duration_ms":83144,"concrete_test":"Build a gold-standard corpus for 2014–2023 by merging the reference lists of prior WAT surveys (Doğan et al. 2014; Balsam & Mishra 2024; the present survey's References) and running a broad search across the five digital libraries plus arXiv/Semantic Scholar using abstract/keyword queries with expanded terms ('web', 'AJAX', 'DOM', 'Selenium', 'JavaScript', 'browser', 'GUI', 'web app', 'web application') and no open-access or title restriction. Compute the fraction of gold-standard papers that appear in the 314-paper corpus, and recalculate the headline distributions (test objectives in §5.2, generation methods in §5.3, execution methods in §6.2) on the gold standard. If recall is below, say, 70%, or if any proportion shifts by more than 10 percentage points, the corpus is not representative and the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of a comprehensive overview rests on the assumption that the 314 papers retrieved in Section 3.2 are representative of WAT research from 2014 to 2023. That assumption is not validated. The search is title-restricted (Table 1), so relevant papers whose titles do not contain 'web' or 'browser based' together with a test-related term (e.g., papers about AJAX, DOM, Selenium, or JavaScript testing) are systematically missed. The inclusion criteria additionally require open access, which excludes paywalled papers from top venues such as IEEE TSE, ACM TOSEM, ICSE, and FSE; these venues publish influential WAT work. The authors assert in Section 3.2 that they are 'confident that the overall trends we report on are accurate,' but no recall or precision check against a gold-standard set is reported. Because every reported proportion—test objectives (§5.2), generation methods (§5.3), execution modes (§6.2), metrics (§7), tools (§8)—is computed from this corpus, any selection bias propagates directly into the survey's empirical claims. Without a representativeness check, the paper cannot support its claim to provide a fair representation of the field.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a systematic literature survey of web application testing (WAT) research published between January 2014 and December 2023. The authors retrieved 314 open-access papers from ACM, Elsevier, IEEE, Springer, and Wiley using title-restricted keyword queries, then classified them according to six research questions covering publication trends, test-case generation, test execution, evaluation metrics, available tools, and challenges/future work. The paper reports proportional distributions for test objectives, generation methods, execution modes, metrics, and tool categories, and it provides an appendix listing all included studies plus a supplementary repository.","tokens_in":48026,"tokens_out":5400,"duration_ms":54588,"significance":"If the corpus is representative, the survey would be a useful reference for researchers and practitioners: it assembles a large catalog of WAT papers, organizes them into a structured taxonomy, and provides tables of datasets and tools with links. The authors make the underlying data available in a supplementary repository, which is a strength for reproducibility. The paper's empirical claims are arithmetic summaries of manual classifications rather than fitted or predicted quantities, so there is no circularity burden. However, the central claim of comprehensiveness rests on an unvalidated corpus-selection procedure, and there are internal contradictions between the announced scope and the executed scope. These issues are load-bearing for a survey whose main contribution is its claimed fair representation of a decade of research.","major_comments":[{"comment":"The Introduction states that the paper explores WAT 'from the perspectives of test case generation and execution, failure diagnosis, evaluation and assessment, regression testing, and available tools,' but Section 3.2 explicitly excludes 'failure diagnosis and regression testing for web applications' from the survey. This is a direct contradiction in the scope of the claimed contribution. Section 7.1 also lists 'failure diagnosis' among primary effectiveness metrics without providing a corresponding analysis subsection. The authors should either include these topics in the corpus and analysis or clearly revise the Introduction and Section 7.1 so that the stated scope matches the executed scope.","section":"§1 and §3.2"},{"comment":"The corpus-selection methodology is not validated against the field, although all reported proportions in Sections 5-8 are computed from this corpus. The search is restricted to titles containing 'web' or 'browser based' together with a test-related term, which systematically excludes influential WAT papers whose titles refer instead to Selenium, AJAX, DOM, JavaScript, or specific vulnerability names. The open-access filter further excludes paywalled papers from venues such as IEEE TSE, ACM TOSEM, ICSE, and FSE. The assertion in Section 3.2 that the authors are 'confident that the overall trends we report on are accurate and provide a fair representation' is unsupported because no recall or precision check against a gold-standard set is reported. I recommend adding a validation study, such as comparing the retrieved set against a manually assembled list of landmark WAT papers, and reporting the resulting recall; alternatively, the paper should be reframed as a survey of open-access, title-matching papers rather than a comprehensive overview.","section":"§3.2 and Table 1"},{"comment":"The description of the snowballing step is inconsistent with Table 2. The table shows 314 papers remaining after applying selection criteria to the keyword-based search results, but the text then says that a snowballing approach 'led to the identification and inclusion of several additional papers' and that 'eventually, 314 papers were included.' If snowballing added papers, the final count should exceed the post-filter count unless the table already includes snowballed papers. This makes the corpus-construction process irreproducible as described. Please clarify the exact role of snowballing and adjust the table or the text accordingly.","section":"§3.2 and Table 2"},{"comment":"The number of research questions is internally inconsistent. The Introduction says the methodology presents 'six research questions,' Section 3.1 lists exactly six RQs, but the Conclusion states that 'Our review was guided by eight research questions.' This inconsistency affects the reader's ability to trust the organization of the survey and should be corrected.","section":"§1, §3.1, and §10"}],"minor_comments":[{"comment":"The pie charts and tables use the Chinese column header '类别 数量' instead of English labels; these should be translated or removed for consistency with the rest of the manuscript.","section":"Figures 4-10"},{"comment":"The sentence 'Sections 4 to ?? address each of the research questions' contains a literal '??' placeholder that should be replaced with specific section numbers.","section":"§1, last paragraph"},{"comment":"The sentence beginning 'The papers analyzed in this study were from various journals and conferences. As shown in, the majority...' has an incomplete cross-reference after 'As shown in'; the figure reference should be completed.","section":"§4.2"},{"comment":"The counts in Figures 8 and 9 sum to 268 and 266, respectively, rather than 314, indicating that some papers report multiple metrics. The text refers to these as proportions 'of the papers' without stating the denominator; please clarify that percentages are computed over metric occurrences, not over papers.","section":"§7.1 and §7.2"},{"comment":"There are citation inconsistencies: LoadRunner is attributed to reference [32] but the LoadRunner discussion appears in reference [87], and the WebQT tool is discussed with reference [73] in Section 7.2.1 while it appears with reference [37] in Section 5.3.3. These citations should be checked and corrected.","section":"§5.2.3 and §7.2.1"},{"comment":"Several cited works are dated 2024, which is outside the stated 2014-2023 publication window for included studies. If these are used only as contextual or future-work references this should be made clear; if they are treated as corpus evidence, they violate the inclusion criteria.","section":"References"},{"comment":"The OWASP AppSensor entry appears twice in the table, and the tool name 'jÄk' is inconsistent with the spelling 'jäk' used in the reference list. These duplicates and typographical inconsistencies should be fixed.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is formatted as a J. ACM submission but appears on arXiv; the front matter (2018 copyright, publication date) needs updating regardless of target venue. The main concern for the editor is that the survey's central contribution is a quantitative map of the field, and that map is built from a corpus whose selection bias is not quantified. This is fixable within the manuscript's scope by adding a validation/recall check or by softening the comprehensiveness claim, but it must be addressed before the paper can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good survey to have on hand, but don't trust the percentages. The paper does something genuinely useful: it organizes 314 open-access papers (2014–2023) into an RQ structure covering test generation, execution, metrics, tools, and challenges, and the supplementary GitHub with the full paper list is a nice touch. The dataset table (Table 3) and tool table (Table 4) are practical references.\n\nThe soft spots are in the methodology, and they matter. The search is title-restricted (Table 1), so papers on AJAX, DOM, Selenium, or JavaScript testing that don't have 'web' or 'browser based' in the title are systematically missed. The open-access filter excludes paywalled work from TSE, TOSEM, ICSE, FSE, which publish influential WAT papers. No recall/precision check is done, and the authors' assurance in Section 3.2 that they're 'confident that the overall trends we report on are accurate' is just an assertion. Since every reported proportion (test objectives, generation methods, execution modes, metrics, tools) is computed from this corpus, the selection bias propagates into the headline numbers.\n\nThere are also internal inconsistencies: the Introduction promises failure diagnosis and regression testing, but Section 3.2 explicitly excludes them; the methodology says six RQs but the Conclusion says eight; cross-references are broken (Sections 4 to ??); and citation numbering is inconsistent (e.g., [1] used for both Selenium and the first dataset). These are fixable but need a careful revision.\n\nOverall: if you want a quick map of WAT subtopics and a list of tools/datasets, this is a decent starting point. But I wouldn't treat the trend percentages as authoritative, and the paper's claim to be 'comprehensive' isn't supported. With a re-run of the search (broader queries, no open-access filter, snowballing on excluded venues) and a validation check against a gold-standard set, this could be a solid survey. As is, it deserves a serious referee but needs major revision before acceptance.","headline":"A useful but flawed survey: the RQ-organized map of 314 papers is handy, yet the unvalidated open-access/title-only corpus undercuts the 'comprehensive' claim.","tokens_in":48549,"tokens_out":1918,"would_cite":false,"duration_ms":19988,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 314 papers maps a decade of web application testing, from test generation and execution to evaluation metrics and tools.","keywords":["web application testing","systematic review","test case generation","test execution","evaluation metrics","testing tools","software quality assurance","vulnerability detection"],"falsifier":"A replication of the search that broadens the inclusion rules (for example, adding closed-access papers, theses, or additional libraries) and finds that the leading shares—vulnerability detection at 29.94% or automated execution at 91.08%—move by more than a few percentage points would demonstrate that the reported landscape is an artifact of the protocol rather than a property of the field.","tokens_in":47601,"feed_emoji":"🧪","tokens_out":7470,"duration_ms":71682,"temperature":0.7,"pith_summary":"This survey sets out to establish a field-level picture of web application testing over the decade 2014-2023, based on 314 primary studies drawn from five major digital libraries. It claims that the literature can be organized around six questions: how the field evolved, how test cases are generated and executed, how results are evaluated, what tools exist, and what challenges remain. If the picture is right, it gives researchers and practitioners a reliable baseline for what has been tried, what has worked, and where the gaps are. The paper also identifies the dominant concerns of the decade, notably security- and vulnerability-oriented testing, and points to scalable automation, standardized metrics, and large language models as the next frontier.","feed_headline":"A 314-paper review maps a decade of web application testing","feed_subtitle":"See which generation methods, execution modes, metrics, and tools dominated web application testing from 2014 to 2023.","key_machinery":"The load-bearing machinery is the literature-review protocol: six research questions, tailored keyword queries over five digital libraries, inclusion and exclusion criteria (English, WAT-related, not theses or prior surveys, open access), and snowballing from reference lists. What carries the argument is the resulting corpus of 314 primary studies, categorized by topic, generation method, execution mode, metric, tool, and challenge; the reported percentages (for example, 29.94% vulnerability detection as an objective, 27.50% model-based generation, 91.08% automated execution) are the concrete evidence for the survey's characterization of the decade.","core_discovery":"The paper's central claim is that its six-question framework, applied to a systematically selected corpus of 314 open-access papers, yields a fair and representative account of a decade of web application testing research. On that basis it reports that web vulnerability detection and security testing together dominate the field's objectives, that model-based testing is the most common test-case generation strategy, that automated execution prevails in local environments, and that security testing tools are the most frequently discussed tool category. It further claims that the main open problems are scaling automation to dynamic and asynchronous applications, maintaining test suites under UI churn, tool fragmentation, and the absence of standardized evaluation metrics.","pith_inferences":["The survey leaves implicit that its corpus, dominated by conference papers and open-access sources, may overrepresent early-announced techniques and underrepresent industrial practice and negative results; a reader should treat the percentages as proportions of the indexed literature, not of actual testing activity.","One consequence the authors do not draw: the heavy concentration on security objectives suggests that web application testing research and web security research are converging, so future surveys may need to treat security tooling as a first-class testing concern rather than a separate discipline.","A testable extension would be to feed the same six-question taxonomy to a corpus that includes theses, closed-access venues, and the 2024-2025 literature to see whether the reported proportions shift; if they shift substantially, the decade picture is an artifact of the search boundary.","A reader should also note an internal inconsistency the paper itself does not address: the conclusion says the review was guided by eight research questions, while Section 3 defines six; the six in Section 3 are the operative ones and the eight appears to be an editing slip."],"forward_implications":["The dominant position of security and vulnerability testing suggests that the field's center of gravity shifted toward attack-surface assurance rather than general functional correctness.","Model-based testing being the most common generation method implies that structural models of application behavior remain the default raw material for automated test creation.","The prevalence of automated execution and local environments indicates that most published work targets repeatable, controlled evaluation rather than distributed, real-world conditions.","The identified lack of standardized metrics means that cross-study comparisons of testing effectiveness and efficiency are not yet trustworthy.","Challenges around dynamic content, asynchronous operations, and UI churn imply that test-suite maintenance, not initial generation, is a key cost driver for web applications."],"supporting_citations":[{"why":"Supplies the systematic review methodology and repository-selection pattern the survey follows.","marker":"[80]"},{"why":"Provides the literature-review approach for analyzing the past to structure the survey.","marker":"[187]"},{"why":"Guides the snowballing procedure and the presentation of findings from a related testing survey.","marker":"[66]"},{"why":"Exemplifies the narrow, layout-focused prior review that the survey positions itself against.","marker":"[145]"},{"why":"An early foundational review of web application testing that this survey extends.","marker":"[104]"},{"why":"A prior systematic literature review of web application testing that defines the field's baseline.","marker":"[50]"},{"why":"A recent review covering only 72 studies, cited as the motivation for a more comprehensive corpus.","marker":"[23]"},{"why":"A focused prior survey on regression test-case generation that shows the need for broader coverage.","marker":"[72]"}],"fun_headline_variants":["314-paper survey maps web testing evolution","Web testing review: security dominates, model-based wins","Decade of web testing: automation meets fragmentation","Survey: web testing gaps in dynamic and async apps","Web testing decade: key findings from 314 studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions depend on the assumption that the 314 papers selected from five digital libraries using title-restricted keyword queries, open-access filtering, and exclusion of theses and prior surveys are a representative sample of web application testing research from 2014 to 2023, an assumption the authors state directly when they say they are confident the overall trends are accurate.","fun_headline_variants_meta":{"raw":{"variants":["314-paper survey maps web testing evolution","Web testing review: security dominates, model-based wins","Decade of web testing: automation meets fragmentation","Survey: web testing gaps in dynamic and async apps","Web testing decade: key findings from 314 studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1273,"prompt_tokens":884,"completion_tokens":389,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":316}},"tokens_in":500,"tokens_out":389,"duration_ms":4487,"temperature":1.0,"reasoning_tokens":316,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:39:55.164456+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication of the search that broadens the inclusion rules (for example, adding closed-access papers, theses, or additional libraries) and finds that the leading shares—vulnerability detection at 29.94% or automated execution at 91.08%—move by more than a few percentage points would demonstrate that the reported landscape is an artifact of the protocol rather than a property of the field.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Exemplifies the narrow, layout-focused prior review that the survey positions itself against."}],"review_version":1}