{"id":"4aec1057-662b-452d-966d-d5e171ceb875","arxiv_id":"2412.03854","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey-based taxonomy that splits reproducibility in AI/ML into eight rigor aspects (repeatability, reproducibility, replicability, adaptability, model selection, label/data quality, meta/incentive, maintainability) with estimated prevalence among 101 papers.","lead":"This paper sorts machine learning research on reproducibility into eight distinct topics, from re-running your own code to keeping systems working over time. It gives the field a shared vocabulary for a debate where the same word currently means many different things.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1 percentages rest on a self-identification filter that is not consistently defined or applied; several Appendix A entries appear not to self-identify as reproducibility papers, so the denominator is uncertain.","rationale":"The reader's weakest assumption correctly identifies corpus representativeness and subjective category assignment as the fragile points. My stress-test agrees but sharpens the issue: the inclusion criterion is not merely possibly unrepresentative, it is internally under-specified. Section 2 says papers 'must self-identify itself as being about repeatability, reproducibility, or replicability,' while Appendix A says papers 'self-identify as being about reproducibility in some sense.' These are different filters, and several included papers on cross-validation and benchmark construction appear unlikely to satisfy the stricter one. If the filter is actually the looser one, then the percentages describe an ad-hoc collection rather than a well-defined population. This is a construct-validity problem, not just a sampling problem. I do not think this changes the verdict from CONDITIONAL: the eight-way taxonomy and the literature summaries are still useful as a proposal, and the paper's own caveats about subjectivity are visible. But the quantitative proportions should not be cited as evidence about the field until the corpus is reconstructed with an explicit protocol and reliability check. Since the reader's verdict already conditions acceptance on exactly these requirements, I leave the verdict unchanged.","tokens_in":23811,"tokens_out":6023,"duration_ms":51771,"concrete_test":"Independent reconstruction: obtain the full text of all 101 Appendix A papers; apply one strict, pre-registered operational rule—only papers whose abstract or introduction explicitly uses 'repeatab*', 'reproducib*', or 'replicab*' as a self-description qualify; then have two annotators blind to the authors' labels assign each qualifying paper to exactly one of the eight categories. Report Cohen's kappa and recompute Table 1. If kappa is below 0.7, or if the qualifying set differs from 101 by more than a few papers, or if any category percentage shifts by more than 5 percentage points, the quantitative proportions in Table 1 are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that Table 1 quantifies the proportion of each rigor type 'as studied today'—depends on two operational choices: (i) a paper enters the denominator only if it self-identifies with 'repeatability/reproducibility/replicability' (Section 2), and (ii) each paper receives one primary category by 'our subjective call' (Appendix A). The first choice is load-bearing because the paper's stated construct is 'scientific rigor,' but the inclusion rule samples papers that merely use the word 'reproducibility.' The second is load-bearing because the categories are not measured, they are assigned. Two concrete problems follow. First, the filter is ambiguous: Section 2 requires the three ACM terms, while Appendix A admits papers that self-identify 'in some sense,' and several listed methodological entries on cross-validation and benchmark construction (e.g., Bates et al. 2021; Bergmeir, Hyndman, and Koo 2018; Varoquaux 2018) do not obviously self-identify with reproducibility. If such papers were included anyway, the denominator is not the population the text describes. Second, with a single rater and no pre-specified rubric for resolving multi-topic papers, the 19.8% Model Selection figure and even the category boundaries are not independently checkable. The paper is honest about subjectivity, but honesty does not turn prevalence estimates into evidence. The taxonomy may still be useful as a proposal; the quantitative contribution is unsupported as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the AI/ML community uses \"reproducibility\" as an umbrella term covering at least eight distinct rigor concerns. It proposes a taxonomy of repeatability, reproducibility, replicability, adaptability, model selection, label/data quality, meta & incentive, and maintainability, defined in Table 1, with the percentage of papers primarily devoted to each aspect reported from the authors' manual categorization of a corpus of 101 papers listed in Appendix A that \"self-identify\" with reproducibility terminology. The paper also surveys the historical literature for each aspect, proposes dependency relations among the rigor types (Section 3, Figure 1), and recommends that major AI/ML conferences create dedicated rigor tracks.","tokens_in":24024,"tokens_out":11153,"duration_ms":106031,"significance":"If the taxonomy is adopted by the community, it would supply a much-needed shared vocabulary and a map of where rigor effort is concentrated, and the paper's historical scholarship — especially the tracing of model-selection and benchmarking concerns back decades in Section 2.4 — is a genuine service to the field. The paper is also commendably transparent: the entire 101-paper corpus is listed in Appendix A, and the authors explicitly disclose that the category assignments are subjective. These strengths make the qualitative contribution real and citable. However, the quantitative claim that the paper has \"quantified the proportion of each type as studied today\" is not supported by the evidence as presented, because the corpus is non-systematic and the inclusion filter is applied inconsistently; the percentages in Table 1 should be treated as an illustrative breakdown of a convenience sample, not as measurements of the field, unless the corpus is rebuilt with a documented protocol or the claim is explicitly demoted.","major_comments":[{"comment":"Section 2 states that a paper qualifies for Table 1 only if it \"must self-identify itself as being about 'repeatability, reproducibility, or replicability',\" but the corpus in Appendix A does not follow that rule: Bates, Hastie, and Tibshirani (2021), Bergmeir, Hyndman, and Koo (2018), Varoquaux (2018), Dacrema, Cremonesi, and Jannach (2019), and Kim et al. (2022) are all categorized as Model Selection papers, yet none of these works frames itself with reproducibility terminology. Appendix A's looser operative criterion (\"self-identify as being about 'reproducibility' in some sense\") is not consistently defined, so the denominator of Table 1 is not the population described in Section 2, and the percentages cannot be taken at face value.","section":"§2 / Appendix A / Table 1"},{"comment":"The corpus is described in Section 2 as \"all literature we are aware of,\" with no search protocol, no statement of venue or database coverage, and no formal inclusion/exclusion criteria beyond self-identification; the Conclusion (§4) nevertheless claims to have \"quantified the proportion of each type as studied today.\" Because the Table 1 percentages (e.g., Model Selection 19.8%, Maintainability 12.9%) are conditional on an unstated and unrepeatable sampling procedure, the quantitative claim is not reproducible by other researchers. The paper should either supply a documented search and screening protocol (keywords, databases, date range, screening steps) or explicitly reframe Table 1 as a descriptive breakdown of the reviewed 101 papers and remove the \"quantified\" claim from the Conclusion.","section":"§2 / Table 1 / §4"},{"comment":"Appendix A discloses that the primary-category assignment is \"our subjective call,\" but no coding rubric is given and no inter-rater reliability is reported; since many of the 101 papers touch multiple aspects, the single-rater assignments determine both the category boundaries and every percentage in Table 1. A concrete example of boundary instability is Musgrave, Belongie, and Lim (2020), coded \"Reproducibility\" in Appendix A but described in §2.2 as a problem of simultaneous, confounding changes to baselines — a concern that fits the paper's own \"Model Selection\" definition at least as well. A dual-coding exercise with reported agreement statistics, or at least a published coding rubric with example papers per category, is needed before the proportions can be treated as measurements.","section":"Appendix A / §2.2"},{"comment":"Even if the inclusion filter were applied consistently, the sampling frame would still not support the Conclusion's framing: the stated construct is the scope of \"scientific rigor\" (Introduction), yet the corpus is restricted to papers that announce themselves with reproducibility terminology, and the paper itself notes that major bodies of relevant work (label-fusion methods in §2.6, decision-tree robustness studies in §2.5) never use such terms. Table 1 therefore systematically undercounts Adaptability and Label/Data Quality, the very categories whose literatures are least self-labeled, so the reported proportions are biased relative to the stated construct; this limitation should be acknowledged explicitly and the claim softened.","section":"Introduction / §2.5 / §2.6 / §4"}],"minor_comments":[{"comment":"The sentence \"the most obvious, and intuitive connections are from repeatability to reproducibility to repeatability\" should read \"...to replicability\"; as written it contradicts the paragraph's own escalation argument (repeatability → reproducibility → replicability).","section":"§3.1"},{"comment":"Section 2.8 contains \"maintinable\" (should be \"maintainable\"), and Figure 1's legend contains \"Eachother\"; more substantively, the dashed line labeled \"Interact With Each Other\" is not explained in the text, so readers cannot tell what distinguishes Maintainability, Data Quality, and Model Selection as an interaction class.","section":"§2.8 / Figure 1"},{"comment":"The Introduction says the 101 papers were \"published since 2017,\" but Appendix A includes Sculley et al. (2015), which violates the stated date window; either update the text or exclude that paper from the corpus.","section":"Introduction / Appendix A"},{"comment":"The discussion of conference questionnaires and guidelines in §2.2 mentions ACM terminology but does not cite the primary ML reproducibility-policy references (e.g., the NeurIPS reproducibility checklist literature), which would strengthen the incentives discussion in §2.7.","section":"§2.2 / §2.7"},{"comment":"The Introduction states that the ACM terminology is \"still insufficient,\" and prior taxonomies are cited (Tatman et al. 2018; Gundersen et al. 2018), but the paper never systematically contrasts its eight categories with these frameworks; a brief comparison would help readers identify the novel contribution and would also address the category instability that Footnote 1 concedes.","section":"§1 / Footnotes 1–2"},{"comment":"Figure 1 distinguishes \"hard dependencies\" (solid lines) from \"influencing effects\" (dashed lines), but the criteria for this distinction are never defined, so the graph reads as a set of intuitions rather than a checkable model; a sentence explaining how a claimed dependency could in principle be tested would help.","section":"§3 / Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The survey/position character of the manuscript sits somewhat awkwardly with the quantitative framing, and the corpus selection is the load-bearing weakness. The first author's own works are well represented in the corpus (e.g., Raff 2019, 2021, 2022; Raff and Holt 2023), which is understandable given his centrality to the area but compounds the difficulty of evaluating the corpus's representativeness; I would welcome a reviewer with survey-methodology experience. That said, the paper is honest about its subjectivity, and the taxonomy itself is useful; the revision path is clear: either rebuild the corpus with a documented protocol or rescope the claims to the reviewed set."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Ed — quick take: this is a genuinely useful synthesis. The eight aspects (repeatability, reproducibility, replicability, adaptability, model selection, label/data quality, meta/incentive, maintainability) are well-chosen and the descriptions are accurate. It goes beyond the ACM trio and prior taxonomies by adding areas like adaptability and maintainability, and the historical-connection narrative is a real contribution. The authors are honest that the category assignments are 'our subjective call,' and the paper reads like a careful map of the literature rather than a sales job.\n\nThe soft spot is exactly where the reader and stress-test point: Table 1's percentages are treated as 'the proportion of each type as studied today,' but the underlying corpus is not systematic. The inclusion rule is described as self-identification with 'repeatability/reproducibility/replicability' in Section 2, yet Appendix A says 'in some sense' and includes several papers that never use those terms—Bates et al. 2021 on cross-validation, Bergmeir et al. 2018, Varoquaux 2018 are methodological papers, not reproducibility self-identifiers. That makes the denominator fuzzy. On top of that, each paper gets one primary category by a single rater with no pre-specified rubric, so the 19.8% Model Selection figure isn't independently checkable. Honesty about subjectivity is good, but it doesn't turn counts into evidence.\n\nThat doesn't sink the paper, because the taxonomy stands on its own as a proposal. The percentages should be reframed as illustrative of the sample, or the authors should provide a reproducible search protocol, inter-rater reliability, and a sensitivity analysis. Either way, the core eight-way distinction is worth having in the literature.\n\nWho's this for? Anyone working on reproducibility in ML, and especially people designing conference tracks or funding calls around rigor. It deserves a serious referee—send it out, but expect the quantitative part to need work. I'd probably cite it for the taxonomy, not for the numbers.","headline":"A useful eight-way taxonomy of rigor in ML, but the prevalence percentages in Table 1 are not robust enough to be read as field measurements.","tokens_in":24614,"tokens_out":2163,"would_cite":true,"duration_ms":76930,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Eight distinct rigor problems hide under the word 'reproducibility'.","keywords":["reproducibility","replicability","scientific rigor","machine learning","taxonomy","literature survey","model selection","maintainability"],"falsifier":"Have several independent coders assign every relevant rigor aspect to each of the same 101 papers instead of one primary category; if many papers touch three or more aspects, or if a broader sample of recent papers falls outside the eight categories, the taxonomy's boundaries and the Table 1 percentages are not stable.","tokens_in":23559,"feed_emoji":"🧪","tokens_out":6305,"duration_ms":61193,"temperature":0.7,"pith_summary":"The paper attempts to establish that the AI/ML community's talk of a 'reproducibility crisis' actually refers to at least eight distinguishable scientific-rigor research programs, all filed under the one word 'reproducibility.' It offers a taxonomy of eight aspects—repeatability, reproducibility, replicability, adaptability, model selection, label/data quality, meta & incentive, and maintainability—and assigns each of 101 self-identifying papers published since 2017 to one primary aspect. The result is a quantitative map of where the community's rigor effort concentrates, with model selection the largest slice at 19.8% and adaptability and label/data quality the smallest at 4.0% each. If the taxonomy holds, the field gains a shared vocabulary and a way to see which rigor problems are well served and which are neglected.","feed_headline":"Eight distinct rigor problems hide under 'reproducibility'","feed_subtitle":"A review of 101 papers shows model choice, data quality, and incentives hiding under one loaded word.","key_machinery":"The load-bearing object is the eight-category taxonomy defined in Table 1, built by a manual review of 101 papers that self-identify as being about repeatability, reproducibility, or replicability. Each category is one named rigor aspect with a one-sentence concern, and the paper uses the taxonomy to do two things: count how often each aspect is the primary focus of current literature, and lay out dependency relations among aspects (solid arrows for hard dependencies such as repeatability before reproducibility, dashed arrows for influences such as model selection and data quality flowing into the other aspects). The taxonomy is what converts an ambiguous slogan into measurable, comparable research topics.","core_discovery":"The paper's central claim is that the community's overloaded use of 'reproducibility' conflates eight distinct aspects of scientific rigor, each with its own question: Can the original authors repeat their own results with their own code and data (repeatability)? Can a different team get the same results using the provided code and data (reproducibility)? Can a different team write new code or use different data and still reach congruent results (replicability)? Can the original method handle new data (adaptability)? How should one reliably choose among competing models (model selection)? How can labeling processes yield stable, low-error labels (label/data quality)? What incentives drive or block rigor (meta & incentive)? And what keeps a solution working as people, code, and data change over time (maintainability)? The paper further claims these aspects interlock: repeatability is a precondition for reproducibility, reproducibility for replicability, and maintainability is essentially iterated replicability over time plus instantaneous repeatability at each point.","pith_inferences":["An implication the authors leave implicit is that funding agencies and journals could use the eight categories as a checklist when asking what rigor aspect a proposal addresses, which would make the neglected categories harder to overlook.","This reader infers that the eight aspects are not cleanly disjoint: most papers touch several, so a multi-label coding of the same corpus would likely shift the percentages and may reveal stable pairs such as reproducibility plus model selection.","A testable extension would be to run the same taxonomy on a broader, automatically sampled corpus across conferences and years to see whether the 4.0% figures for adaptability and label/data quality are a stable property of the field or an artifact of this review."],"forward_implications":["Researchers and reviewers can say which of the eight aspects a given 'reproducibility' study actually addresses, removing a recurring source of confusion.","Because the aspects form a dependency chain, a failure at the repeatability level undermines any downstream reproducibility or replicability claim.","Maintainability can be understood as repeatability plus replicability stretched over time, so aging software, drifting labels, and hardware changes are not separate concerns but the same rigor problem.","The measured proportions identify model selection as the most-studied aspect and expose adaptability and label/data quality as the most neglected, giving a concrete agenda for where new rigor research would matter.","The paper's recommendation to create a dedicated scientific-rigor track at major AI/ML conferences is a direct corollary of having a named set of topics to consolidate."],"supporting_citations":[{"why":"Documents the confused and incompatible terminology that motivates the paper's disentangling effort.","marker":"(Plesser 2018)"},{"why":"Provides the state-of-the-art survey of AI reproducibility that anchors the corpus of self-identifying papers.","marker":"(Gundersen and Kjensmo 2018)"},{"why":"Supplies the only large-scale code-execution baseline the paper cites, reporting that 74% of released code ran without issue.","marker":"(Trisovic et al. 2022)"},{"why":"Underpins the replicability category with a 255-paper reimplementation study and features that predict replicability.","marker":"(Raff 2019)"},{"why":"Establishes the maintainability rigor aspect through the concept of hidden technical debt in ML systems.","marker":"(Sculley et al. 2015)"},{"why":"Provides the canonical example of unsound comparisons that motivates the reproducibility-in-depth and model-selection categories.","marker":"(Musgrave, Belongie, and Lim 2020)"}],"fun_headline_variants":["Reproducibility: one word, eight distinct rigor issues","Eight rigor facets hidden in one loaded term","The many meanings of 'reproducible' in ML research","When 'reproducible' means eight different things"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative map stands on the assumption that the 101 manually collected, self-identifying papers fairly represent the field and that each paper can be assigned one primary rigor category.","fun_headline_variants_meta":{"raw":{"variants":["Reproducibility: one word, eight distinct rigor issues","Eight rigor facets hidden in one loaded term","The many meanings of 'reproducible' in ML research","When 'reproducible' means eight different things"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1397,"prompt_tokens":847,"completion_tokens":550,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":485}},"tokens_in":463,"tokens_out":550,"duration_ms":5297,"temperature":1.0,"reasoning_tokens":485,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:59:49.747833+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have several independent coders assign every relevant rigor aspect to each of the same 101 papers instead of one primary category; if many papers touch three or more aspects, or if a broader sample of recent papers falls outside the eight categories, the taxonomy's boundaries and the Table 1 percentages are not stable.","supporting_citations":[],"review_version":1}