{"id":"2cb1ce7a-af40-4baf-a3c0-60467a081ba7","arxiv_id":"2501.15662","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A retrospective argument that publishing a baseline with data and scripts created a research community, and that the old PROMISE datasets now slow the field down.","lead":"In this retrospective, the original author of the 2007 NASA defect-prediction paper argues that publishing a baseline with data and scripts can seed a whole research community. He also argues that the field has become too dependent on stale datasets and should move to fresher data and methods.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central recipe—publish a baseline plus data/scripts and a community will form—is inferred from a single self-authored case, with venue, first-mover advantage, and PROMISE conference/repository scaffolding all confounded; a matched historical comparison is needed.","rationale":"The paper is best read as a retrospective essay with a strong, generalizable causal thesis: that open sharing of a baseline plus data and scripts is sufficient to create a research community. The essay's own historical account, however, bundles that act together with several other deliberate community-building interventions: a dedicated conference series, a large curated repository, active solicitation of datasets, and the recruitment of prestigious steering-committee members. It also relies on a single case in which the author was the central actor. This is not an accusation of bad faith; it is a standard identification problem. The quantitative claims about citations and artifact adoption may well be correct, but they support correlation, not the stronger 'just by publishing' causal claim. The reader's weakest-assumption analysis already identifies the same confound, so I agree with that assessment. The appropriate verdict remains UNVERDICTED: the essay is valuable and plausible, but its central causal claim is not established by the supplied evidence. The concrete test proposed above would convert the retrospective into a testable historical claim.","tokens_in":11480,"tokens_out":4202,"duration_ms":41104,"concrete_test":"Assemble a cohort of all defect-prediction papers published in TSE/ICSE/FSE between 2003 and 2010 (target n≈40-60). For each paper, code: (a) public availability of data and scripts, (b) venue tier, (c) first-baseline vs follow-up status, (d) author prior citation impact, and (e) whether an accompanying conference/repository or outreach program existed. Measure five-year community formation: cumulative citations, count of external papers reusing the dataset, and derivative datasets created. Then estimate the association between (a) and community formation, with and without controls for (b)-(e), or run a matched comparison where artifact-sharing papers are paired with non-sharing papers matched on (b)-(e).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central assertion is causal and general: 'Industry can get any research it wants, just by publishing a baseline result along with the data and scripts needed to reproduce that work' (Abstract; echoed in Conclusion). The 2007 paper is offered as the proof. But the evidence bundle in Section 1 includes far more than the shared artifact. The PROMISE project supplied an annual conference, a repository of hundreds of datasets, weekly student sprints that solicited reproduction data from authors, and a steering committee that 'earned the prestige needed for future growth.' The 2007 paper was also the first public defect-prediction baseline published in a top venue (TSE), in a subfield that was already hungry for data. None of these confounds—venue, first-mover status, and community scaffolding—is controlled for. The reader's weakest assumption is precisely this causal attribution: the observed popularity of the 2007 paper and PROMISE data is credited to open sharing, but no counterfactual, matched comparison, or competing-explanation test is presented. Since the author is both the creator of PROMISE and the promoter of the lesson, the n=1 case is also self-referential: it cannot by itself establish that a bare baseline-plus-data release is sufficient to create a community. The paper's quantitative facts (citation counts, 20% adoption by 2018) are plausibly accurate, but they establish correlation, not the 'just by publishing' mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This arXiv note is a retrospective on Menzies, Greenwald, and Frank's 2007 IEEE TSE paper on data-mining static code attributes for defect prediction. The author, also the originator of the PROMISE project that hosted the data and scripts, argues that the 2007 paper's open sharing of a baseline result, data, and scripts catalyzed a research community and that industry can deliberately create such communities by publishing baselines with reproducible artifacts. The note reviews the paper's technical claims, reports progress in defect prediction, proposes a four-phase model of shared-data lifecycles (from resistance to stagnation), lists nine 'Menzies's Laws' about software-engineering data mining, and closes with future directions, including changes to his own editorial policy at the Automated Software Engineering journal. The manuscript is written as a personal, opinionated essay rather than a systematic empirical study.","tokens_in":11738,"tokens_out":3091,"duration_ms":30392,"significance":"If the central causal claim—that publishing a baseline plus data/scripts is sufficient to create a productive research community—were established, it would be a practically important and broadly applicable recipe for accelerating research. The retrospective also documents a historically influential dataset collection (PROMISE) and a widely cited paper, and it makes a falsifiable prediction that the same mechanism can be reused for new topics. The paper's strengths include naming a concrete case, providing some external indicators of impact (citation counts, an alleged 20% adoption in leading TSE papers), and being transparent about the author's own role. However, the evidence presented is largely anecdotal and self-referential: the causal claim rests on a single case study whose confounds are not examined, the 20% statistic is not accompanied by a reproducible methodology, and the nine laws are generalizations from the author's own experience and self-cited studies. As a personal retrospective, the narrative is coherent and readable, but as a scientific argument for the universal mechanism, it is not yet supported.","major_comments":[{"comment":"The central claim, 'Industry can get any research it wants, just by publishing a baseline result along with the data and scripts needed to reproduce that work,' is causal and general, but it is supported only by a single case study with several confounds: the 2007 paper appeared in a top journal (TSE), was the first public defect-prediction baseline in that venue, benefited from the PROMISE conference and repository infrastructure (including student sprints and a steering committee), and was promoted by the author himself. The manuscript does not consider or control for these alternative explanations, nor does it offer any counterfactual or matched comparison. Please either temper the claim to a hypothesis or personal observation, or add a systematic comparison (e.g., similar papers that shared data but did not create communities, or communities formed without shared baselines).","section":"Section 1"},{"comment":"The assertion that 'By 2018, twenty percent of leading TSE papers (according to Google Scholar Metrics), incorporated artifacts introduced and disseminated by this research' is load-bearing for the impact narrative, but no methodology is provided. How were 'leading TSE papers' selected and counted? What counts as 'incorporated artifacts'—citing PROMISE data, using the 2007 data, or using any data from the PROMISE repository? Is the count reproducible? Without a precise definition and derivation, this statistic cannot be verified and should not be presented as an established fact.","section":"Section 2"},{"comment":"The nine 'Menzies's Laws' are presented as general empirical laws, but they are each supported by anecdotal evidence, often from the author's own papers. For example, 'Menzies's 6th Law: Data quality matters less than you think' is based on a single mutation study described in the text, and 'Menzies's 5th Law: Bigger is not necessarily better' is inferred from a systematic review that only 13/229 LLM papers compared to other methods—which does not itself demonstrate that smaller models are better. These statements are better framed as personal reflections or working hypotheses, with explicit caveats about the limited evidence, rather than as laws.","section":"Section 5.2"},{"comment":"The four-phase model of shared-data lifecycles ('Data? Good luck with that!' through 'A graveyard of progress') is presented as a general pattern, but it is based on the PROMISE experience alone. The claim that PROMISE data eventually became a 'lead weight' that stifled research is not supported by evidence; the examples of papers reusing decades-old datasets could be interpreted as evidence of continued utility, not stagnation. The author's own editorial decision to desk-reject papers using his 2005 datasets does not establish a field-wide problem. Please either present this as a personal narrative or provide systematic evidence of the claimed stagnation and its causes.","section":"Section 5.1"}],"minor_comments":[{"comment":"The manuscript contains numerous typos and grammatical errors that should be corrected: 'scripts need to reproduce' (Abstract), 'Those result were' (Abstract), 'halycon' (Section 1), 'wore no suite and tie' (Section 1), 'gather all can that be collected' (Section 2), and 'more one attribute' (Section 2).","section":"Throughout"},{"comment":"The statement that the PROMISE repository 'grew so large and that it we moved it to the Large Hadron Collider' is confusing and appears to be a joke or error; if it is a reference to the 'Seacraft' data at Zenodo, clarify the wording.","section":"Section 1"},{"comment":"Several references are self-citations (e.g., [24] duplicates [1]; [7], [23], [34], [39], [43], [44], [50], [57], [59], [61], [62], [64], [65], [66], [69], [70], [71], [72], [73], [77], [80], [85], [86], [87], [88]), which is understandable in a retrospective but should be presented carefully so that the independence of the evidence is clear.","section":"References"},{"comment":"The 'Menzies's 3rd Law: Turkish toasters can predict for errors in deep space satellites' is stated without a reference to the transfer-learning study [67] in the surrounding text, making it hard to evaluate; add a citation and a brief explanation of the study design.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"This is a retrospective/opinion piece rather than a primary research contribution. If the venue is a magazine or perspectives venue, the narrative style and personal voice are acceptable, but the central causal claim and the 'laws' go beyond personal reflection and require substantial reframing or additional evidence. The manuscript is also heavily self-referential: the author is the subject of the case study, the creator of the PROMISE infrastructure cited as the mechanism, and the evaluator of the impact (as editor of a journal that now desk-rejects papers using his own datasets). This does not disqualify the piece, but it strengthens the need for external evidence or explicit caveats about alternative explanations. The 20% statistic and the nine laws, if retained as factual claims, need to be either replaced with peer-reviewed evidence or clearly identified as personal opinions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, read this as a memoir rather than a research paper. It is Tim Menzies looking back at his 2007 TSE defect-prediction paper and the PROMISE repository he built around it. The real content is in the later sections: the four-phase lifecycle of shared datasets (gold standard to graveyard), the announcement that as ASE editor he now desk-rejects papers built on his own 2005 datasets, and nine \"Menzies laws\" distilled from two decades of work. Some laws are well supported by prior results (e.g., data quality matters less than expected, from Shepperd et al.'s mutation experiments); others are provocative but anecdotal (Turkish toasters predicting satellite errors). The voice is candid and the historical detail is genuinely useful, especially for students learning how research communities form and stagnate.\n\nThe soft spot is the framing claim: \"Industry can get any research it wants, just by publishing a baseline result along with the data and scripts.\" That is a causal hypothesis inferred from one highly successful case, and the case is full of confounds: the 2007 paper was the first public baseline in a top venue, and PROMISE came with a conference, a repository, and a steering committee that actively solicited and curated data. The observed popularity is consistent with open sharing, but nothing here tests that against alternatives. The 20% adoption figure is cited without a reproducible method. As memoir, none of that is disqualifying; as science, the central claim is unproven.\n\nThis paper contains no new technical result and should not be reviewed as one. But it is a serious, readable perspective that can inform editorial policy and research practice. I would send it to peer review as an experience/history piece, with a request to soften the causal language and add explicit caveats about confounds. It would spark a good reading-group discussion on open science.\n\nRecommendation: send to review, expect revision.","headline":"A candid first-person history of the PROMISE community, with a plausible but untested causal claim; review it as a perspective piece, not a research paper.","tokens_in":12245,"tokens_out":1981,"would_cite":false,"duration_ms":19651,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This retrospective claims that publishing a baseline result with its data and scripts is enough to grow a research community, and uses a 2007 defect-prediction paper as the case study.","keywords":["defect prediction","reproducible research","data sharing","software engineering","static code metrics","open science","baseline result","dataset reuse"],"falsifier":"One could compare two groups of papers on the same new topic, one published with full data and scripts and one without, matched for venue and author promotion, then count citations, replications, and derivative studies after the same time window; if shared-artifact papers do not grow faster, the central causal claim fails.","tokens_in":126,"feed_emoji":"📊","tokens_out":4717,"duration_ms":94725,"temperature":0.7,"pith_summary":"This retrospective argues that a research community can be created at will: publish a solid baseline result together with the data and scripts needed to recreate it, and other researchers will take up the result, extend it, and build a field around it. The case study is a 2007 software-defect-prediction paper that shared exactly that, and whose shared artifacts were later used in a fifth of leading software-engineering papers. The essay also warns that the same shared resource can outlive its usefulness, turning from a gold standard into a constraint on progress. A sympathetic reader should take away a recipe for building open research communities and a caution about dataset stagnation.","feed_headline":"Publish the data, and a research field follows","feed_subtitle":"A retrospective argues that reproducible baselines can seed any research community—then warns shared data can become a trap.","key_machinery":"The mechanism is the publish-the-baseline loop: an author produces a first credible result, releases the dataset and all scripts needed to reproduce it, and invites others to beat it. The named object in the case study is the shared artifact repository and its companion conference, which supplied hundreds of software-engineering datasets and the scripts that made re-analysis cheap. This loop does the work by lowering the cost of entry for other researchers, turning a single paper into a standard benchmark, and creating a visible target that the field can collectively improve on.","core_discovery":"The central discovery claimed is that open artifact sharing—not just the intellectual contribution of the model—was the load-bearing ingredient that turned a single defect-prediction baseline into a sustained, reproducible research movement. According to the retrospective, the 2007 paper's decision to release its data-mining scripts and datasets allowed hundreds of later studies to compare against, reuse, and improve the baseline; at its peak in 2016 the paper was software engineering's most cited paper per month, and by 2018 a fifth of leading journal articles used data introduced by that line of work. The paper further asserts that the same mechanism can be repeated for any research topic: a community forms when a baseline plus reproduction scripts are published. It also records a set of empirical observations about that body of work, including that different projects favour different metrics, static code attributes matter only in combination, much of the collected data can be discarded without hurting predictions, data-quality flaws matter less than expected, weak learners can still rank solutions usefully, and many supposedly hard problems yield to simpler methods.","pith_inferences":["If the causal recipe is right, research fields could deliberately seed new topics by publishing baselines; a testable extension is for new fields to try the same recipe and measure citation or reuse growth against comparable papers without shared artifacts.","The four-phase trajectory of shared repositories implies that such resources need governance—sunsetting, refresh, or quality re-certification—something the paper only gestures at.","The transfer-learning observation, where a model learned on one industry's data predicts failures in an unrelated system, hints that software metrics live on a low-dimensional structure; if so, data-efficient methods should keep outperforming brute-force data collection on tabular software-engineering tasks."],"forward_implications":["Publishing a reproducible baseline can seed a new research community on any topic, not just defect prediction.","A shared dataset quickly becomes the standard benchmark, which accelerates early growth but later risks locking a field into outdated data.","Defect prediction built on static code attributes is viable and has moved into industry, where surveys show most practitioners willing to adopt it and case studies report reduced inspection effort.","Methodologically, simple comparisons and cheap baselines should benchmark sophisticated methods; in this body of work simpler approaches repeatedly matched or beat complex ones.","Data quality, within limits, is less important than the predictive signal: injecting known quality issues into datasets did not degrade learned-model performance."],"supporting_citations":[{"why":"Supplies the 2007 baseline result and shared scripts that are the case study for the whole argument.","marker":"[1]"},{"why":"Documents that over 90% of surveyed practitioners were willing to adopt defect prediction, evidence of downstream industry acceptance.","marker":"[19]"},{"why":"Commercial case study reporting defects predicted and inspection effort reduced, evidence that the line of work paid off.","marker":"[11]"},{"why":"Industrial deployment of a defect-prediction model at a large electronics company, evidence that the approach works in practice.","marker":"[20]"},{"why":"Comparison showing statistical defect prediction can match static bug finders in cost-effectiveness, supporting the adaptability claim.","marker":"[18]"},{"why":"Audit identifying quality problems in the shared datasets, grounding the claim that data quality matters less than expected.","marker":"[89]"},{"why":"Systematic review showing few large-language-model papers compared against other methods, grounding the critique that complex methods need simple baselines.","marker":"[81]"},{"why":"Demonstrates that poorly predicting models can still rank solutions, grounding the observation about bad learners making good conclusions.","marker":"[90]"}],"fun_headline_variants":["Open data built a whole research field","Reproducible baselines seed entire communities","The key to a cited paper: share your data","Data sharing, not models, drove a research boom","What made a defect-prediction paper iconic? Its data"],"cache_read_input_tokens":14336,"weakest_assumption_plain":"The whole story assumes that the shared data and scripts, rather than being first, the venue's prestige, or the author's promotion, caused the paper's popularity and the community's growth; no comparison group or counterfactual is tested.","fun_headline_variants_meta":{"raw":{"variants":["Open data built a whole research field","Reproducible baselines seed entire communities","The key to a cited paper: share your data","Data sharing, not models, drove a research boom","What made a defect-prediction paper iconic? Its data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1190,"prompt_tokens":865,"completion_tokens":325,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":251}},"tokens_in":481,"tokens_out":325,"duration_ms":3494,"temperature":1.0,"reasoning_tokens":251,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:03:01.753340+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One could compare two groups of papers on the same new topic, one published with full data and scripts and one without, matched for venue and author promotion, then count citations, replications, and derivative studies after the same time window; if shared-artifact papers do not grow faster, the central causal claim fails.","supporting_citations":[{"cited_title":"Percep- tions, expectations, & challenges in defect prediction,","cited_arxiv_id":null,"evidence_quote":"Documents that over 90% of surveyed practitioners were willing to adopt defect prediction, evidence of downstream industry acceptance."},{"cited_title":"Remi: defect prediction for efficient api testing,","cited_arxiv_id":null,"evidence_quote":"Industrial deployment of a defect-prediction model at a large electronics company, evidence that the approach works in practice."},{"cited_title":"Comparing static bug finders and statistical prediction,","cited_arxiv_id":null,"evidence_quote":"Comparison showing statistical defect prediction can match static bug finders in cost-effectiveness, supporting the adaptability claim."},{"cited_title":"Data quality: Some comments on the nasa software defect datasets,","cited_arxiv_id":null,"evidence_quote":"Audit identifying quality problems in the shared datasets, grounding the claim that data quality matters less than expected."},{"cited_title":"Using bad learners to find good configurations,","cited_arxiv_id":null,"evidence_quote":"Demonstrates that poorly predicting models can still rank solutions, grounding the observation about bad learners making good conclusions."}],"review_version":1}