{"id":"fa9882b6-da14-40e0-b326-6247e3e75b09","arxiv_id":"2608.09531","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The HEP-ML Living Review is frozen and replaced by a curated, annotated Living Guide designed for orientation in a mature field.","lead":"The authors of the HEP-ML Living Review are freezing their exhaustive bibliography and replacing it with a curated, annotated guide. They argue that particle physics machine learning has grown too large for complete lists to be useful, so newcomers need structured entry points.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Guide's advantage over the archived comprehensive Review is asserted, not measured: Section 2.3's claim that newcomers want a few curated papers rests on undocumented conversations. A controlled user study would settle it.","rationale":"The paper's strongest claim is that the field has outgrown near-complete manual bibliographies and that a curated Guide is the right replacement. The first half is well supported by the INSPIRE-HEP growth figure and by the described maintenance burden. The second half—that the Guide serves users better—is the part that actually justifies the decision to freeze the Living Review. That half depends on undocumented informal conversations in Section 2.3. This is not an internal inconsistency or a disagreement with consensus; it is a missing measurement in an otherwise transparent infrastructure paper. The resource itself is real and publicly accessible, the archived Review is preserved with a DOI, and the governance model is explicit, so the downside risk is limited. Still, the central prescriptive claim is a design hypothesis until tested. The reader's verdict of CONDITIONAL already captures this, so I do not change it; the paper should either add a user-needs survey or usage study, or explicitly frame the Guide as a testable hypothesis rather than a demonstrated improvement.","tokens_in":6225,"tokens_out":3627,"duration_ms":36493,"concrete_test":"Run a pre-registered controlled user study: recruit 60 target users (e.g., graduate students and postdocs at HEP-ML events) and randomly assign 30 to the archived Living Review plus INSPIRE search and 30 to the new Living Guide. Give each an unfamiliar subfield (e.g., unfolding with ML, equivariant networks, hardware-aware triggering) and ask them to (a) identify the 3-5 foundational papers and (b) plan an entry reading path. Measure time to first accurate answer, completeness against expert labels, and self-reported confidence. If the Guide group is not at least as fast or accurate, the 'particularly useful' claim fails and the freeze decision is not justified by evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The factual premise of the paper is solid: Figure 1 documents the growth in HEP-ML literature, and the citation counts in Section 2.3 show real community use of the original Living Review. The questionable step is the prescriptive move from 'comprehensive lists are no longer enough' to 'a curated Guide is more useful.' Section 2.3 ('What users need now') supports this move with 'Conversations with researchers at all career stages show a consistent pattern,' but no sampling frame, protocol, or results are given. The design in Section 3.2 (three to six starting points, annotated curated lists, no completeness) is exactly right for that assumed need, and exactly wrong if many researchers use near-comprehensive lists systematically, e.g., for literature reviews, reference chasing, or citation audits. Those users are not served by the new Guide; they are displaced to the frozen archive plus search tools. The paper's own evidence shows the Living Review is actively cited (81 entries citing only the URL, 217 only the arXiv article, 20 citing both), indicating real use of comprehensive lists, though not necessarily for navigation. Because the authors' 'better' claim is a design hypothesis, not a demonstrated result, it is the load-bearing assumption. The sustainability prong is also asserted rather than measured, but even a sustainable Guide is not an improvement unless it actually helps the claimed users.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper announces a directional change in a widely used community resource for machine learning in particle physics (HEP-ML). The authors argue that the HEP–ML Living Review, a near-comprehensive community-maintained bibliography, is no longer sustainable or as useful as it was in 2020, because the literature has grown by more than an order of magnitude and has diversified into many subfields. They therefore freeze the Living Review as an archival snapshot dated 1 June 2026 and replace it with a curated HEP–ML Living Guide, which offers annotated recommended starting points, curated paper lists, benchmarks, and software references organized along two independent axes (HEP application and ML method). The paper describes the history and lessons of the original Review, presents publication-growth data from INSPIRE-HEP, reports citation statistics for the Review, and details the structure, scope, and maintenance model of the new Guide.","tokens_in":6407,"tokens_out":3648,"duration_ms":33993,"significance":"The paper is a practical community-resource transition statement rather than a technical research contribution. Its factual premises are documented: Section 2.2 and Figure 1 provide INSPIRE-HEP query results showing the growth of HEP-ML literature, and Section 2.3 gives concrete INSPIRE citation counts (81 URL-only, 217 arXiv-only, 20 both) for the Living Review. The authors are transparent about what they did and did not measure, and they explicitly acknowledge the self-referential nature of evaluating their own project. The significance, if the new Guide succeeds, is real but modest: it could improve onboarding and orientation for newcomers and cross-subfield researchers. However, the central prescriptive claim—that a curated guide is better for users than a near-comprehensive bibliography—rests on anecdotal evidence. The paper does not currently provide enough empirical support to justify the strength of the conclusion, and the sustainability argument is asserted rather than demonstrated. Still, the direction is defensible and the proposed resource is a reasonable community experiment; the claims need tempering or additional evidence.","major_comments":[{"comment":"The claim that students and postdocs want 'three to five papers to read first, two or three tutorials or reviews, and a map of how the subfield connects to its neighbors' is supported only by 'Conversations with researchers at all career stages,' with no number of conversations, no sampling frame, no protocol, and no results. This is the load-bearing empirical premise for abandoning near-comprehensive coverage and moving to a curated guide. As written, it is a design hypothesis, not a finding. To make the argument rigorous, the authors should either (a) report a user study or survey with a described methodology and results, (b) document the anecdotal evidence in a more falsifiable way (e.g., number of researchers consulted, their career stages, the questions asked, and the range of responses), or (c) explicitly reframe the premise as a design assumption and soften the conclusion accordingly. As it stands, the paper's central move from 'comprehensive lists are no longer enough' to 'a curated Guide is more useful' is not fully supported by the presented evidence.","section":"Section 4, Conclusions"},{"comment":"The conclusion states that the field 'has become too large and too mature for a near-complete manual bibliography to stay sustainable or particularly useful.' The 'particularly useful' clause is stronger than the evidence in Section 2.3 supports: the INSPIRE search returns 81 entries citing only the URL, 217 citing only the arXiv article, and 20 citing both, which indicates ongoing use of the comprehensive resource even if the exact usage mode (navigation vs. bibliographic reference) is not known. The new Guide's design rationale is about orientation, so the more precise claim would be that a near-complete list is less useful for early orientation and navigation, while still acknowledging the archival value for comprehensive searches, reference chasing, and citation audits. The current wording overstates the case and should be revised to distinguish these functions.","section":"Section 3, Commitment 5"},{"comment":"The commitment to 'Sustainability by design' is asserted but not demonstrated. The paper explains why the old model was unsustainable at the scale of several hundred new papers per year, but the new Guide's maintenance model depends on named community members writing sections as one-time contributions, with sections staying static until someone updates them. The paper does not address the risk that some subfields, especially smaller or less active ones, may never attract a section author, which would undermine the Guide's goal of providing a map of subfield connections. The self-selection mechanism ('contributions follow real community investment') may produce uneven coverage biased toward currently popular areas. The authors should discuss a fallback plan for orphaned sections, perhaps by listing a provisional roster of section editors or by describing how the coordinators will handle gaps.","section":"Section 3, Commitment 5"}],"minor_comments":[{"comment":"The text says 'It now adds that many in a few months,' referring to roughly 200 papers per year in the first year. The claim would be clearer with a specific number or a comparison of the slopes in Figure 1, for example, the number of papers added in the most recent year versus the first year.","section":"Section 3.2"},{"comment":"The two-axis taxonomy (HEP application and ML method) is described as the organizing principle, but the mechanics of cross-linking are not specified. It would be helpful to state how cross-links are represented in the version-controlled repository (e.g., tags, a lookup table, or bidirectional links in the markdown) so that contributors know how to maintain them.","section":"Section 3.2"},{"comment":"The sentence 'We do not expect this to trigger controversial discussions that would call for extra layers of moderation' presupposes an outcome. If a contested contribution does arise, the governance model does not specify a resolution process. A one-sentence contingency statement would make the maintenance model more robust.","section":"Section 3.2"},{"comment":"The paper's title, 'The Living Guide of Machine Learning for Particle Physics,' is identical to the name of the new resource. A subtitle such as 'A Transition from the Living Review' would make the article's purpose immediately clear to readers searching for the old Review. Also, Reference [1] (the arXiv paper for the original Review) has no DOI while Reference [3] has a Zenodo DOI; please verify that the citation details are consistent with the description in Section 3.2 of the archived Resource.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a community-resource transition announcement, not a conventional research paper. The factual documentation (INSPIRE queries, citation counts, growth figures) is adequate. The main risk is the unsupported prescriptive claim about user needs, which the skeptic's note correctly identifies. If the journal is willing to publish this as a 'community resource' paper, the authors should either add a small user survey or substantially soften the language from 'what users need' to 'our design assumption.' The sustainability risk (Section 3, Commitment 5) is worth a careful response. Overall, the paper is likely publishable after revision, but I would not accept it in its current form because the central claim about usefulness is stronger than the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a clear, honest announcement of a real community resource, but its central justification is a design hypothesis rather than a measured result. What is new: the authors freeze the HEP-ML Living Review and introduce the Living Guide, a curated, annotated, two-axis field guide with named section authors, timestamps, and Zenodo DOI releases. That is a genuinely useful piece of infrastructure, and the paper gives a coherent rationale for it.\n\nWhat the paper does well: it documents the growth of the literature with an INSPIRE-HEP query (Figure 1), reports citation counts showing the old review was used, and is transparent about the limits of its own evidence. It also preserves the old review as a frozen archive, so nobody loses the comprehensive record.\n\nThe soft spot is exactly where the stress-test points. Section 2.3's 'What users need now' rests on 'conversations with researchers' with no sampling frame or protocol. The claim that newcomers want three to five papers rather than a comprehensive list is plausible but untested. The design of the Guide follows from that assumption, so the whole transition leans on it. The authors could fix this by reframing the Guide as a hypothesis and committing to a modest user study. As written, the paper overstates the certainty of the 'better serves users' claim.\n\nI do not think this is a fatal flaw. The resource is low-risk and complementary; the archive remains. But the paper's central rhetorical move, 'the old model no longer serves this field well,' is an assertion, not a demonstrated conclusion. A careful referee should push on that.\n\nThe paper is not a physics result, and it does not pretend to be. It is for community infrastructure and onboarding. I would cite it if I wrote a HEP-ML paper and wanted to point to the Guide. I would bring it to a reading group only if the group cares about scholarly communication.\n\nRecommendation: send it to peer review with a request to explicitly treat the user-needs claim as a design hypothesis, and to add a short reflection on how the Guide's effectiveness will be measured. That would make it a stronger paper.","headline":"A well-documented announcement of a useful new community resource whose core 'curation is better' claim is plausible but unmeasured.","tokens_in":6988,"tokens_out":2676,"would_cite":true,"duration_ms":24871,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The HEP-ML Living Review is frozen and succeeded by a curated, annotated Living Guide because near-comprehensive manual bibliographies have become unsustainable and less useful as the field has grown by an order of magnitude.","keywords":["machine learning","particle physics","living review","curated guide","scientific literature","community curation","HEP-ML"],"falsifier":"Count citations and contributions to the archived Living Review versus the Living Guide over the next two years: if the frozen review continues to be cited at its pre-freeze rate while the Guide attracts few contributors or readers, the claim that near-complete bibliographies are no longer particularly useful would be falsified.","tokens_in":5971,"feed_emoji":"🧭","tokens_out":11381,"duration_ms":88951,"temperature":0.7,"pith_summary":"This paper announces the end of one model of scientific bibliography and the start of another. The authors argue that the HEP-ML Living Review, a community-maintained list of more than 4,000 machine-learning papers in particle physics, has become too large and too mature for near-comprehensive manual curation to remain sustainable or particularly useful. They freeze the review as an archival record covering the literature up to 1 June 2026 and replace it with the HEP-ML Living Guide, a curated field guide that recommends starting points, annotates foundational and representative work, and cross-links applications with methods. The argument matters because it diagnoses a general shift in how a fast-growing field organizes its own literature, from exhaustive accumulation to curated navigation.","feed_headline":"HEP-ML review freezes as field outgrows exhaustive lists","feed_subtitle":"Annotated starting points and a subfield map replace the 4,000-paper list.","key_machinery":"The central mechanism is the HEP-ML Living Guide itself: a versioned field guide organized along two independent axes (HEP application and ML method) with cross-links between them. Every included paper or cluster carries a short annotation stating why it was chosen and what it establishes. Sections are written by named contributors, timestamped, and periodically released as citable snapshots with their own DOI, so a citation points at a fixed state and at the people who produced it. The selection criteria are foundational importance, methodological clarity, and practical use for newcomers, and the resource deliberately makes no claim of completeness.","core_discovery":"The central claim is that the near-comprehensive, community-contributed Living Review model that served machine learning in particle physics from 2020 has been overtaken by the field it helped document. The literature grew from roughly two hundred papers a year to several hundred a year, the flat taxonomy could no longer answer questions like which papers established simulation-based inference or which architectures work for fast simulation, and the community built its own ecosystem of reviews, benchmarks, and software. The paper therefore freezes the Living Review as an archival snapshot up to 1 June 2026 and replaces it with the HEP-ML Living Guide, a curated field guide that does not claim completeness and instead recommends, annotates, and maps the literature.","pith_inferences":["The two-axis taxonomy and annotation requirement could become a template for other disciplines facing literature explosion, making this paper a reusable model rather than a one-off change.","By leaving out nuclear, heavy-ion, and astroparticle physics, the Guide may push those communities to maintain their own linked guides, leading to a federated network of curated entry points.","The authors' informal user interviews could be replaced by a systematic usage study; one testable prediction is that newcomers reach a working understanding of a subfield faster through recommended starting points than through comprehensive lists.","Because section coverage follows community interest, well-funded subfields may receive frequent updates while smaller areas stagnate, a silent bias the paper acknowledges only indirectly."],"forward_implications":["Newcomers to any subfield can start from a short list of recommended papers and tutorials instead of a multi-thousand-entry bibliography.","Citations to community resources become tied to a fixed version and to the named authors of each section, repairing the credit and reproducibility problems of a rolling document.","The frozen Living Review remains available as an archival snapshot, preserving near-comprehensive coverage of the literature up to 1 June 2026 for historical citation.","The Guide deliberately narrows its scope to particle physics, linking to neighboring curated resources for cosmology, astroparticle physics, and accelerators rather than duplicating them.","Maintenance becomes sustainable because named experts write sections as one-time contributions, sections stay until someone updates them, and releases are tagged periodically."],"supporting_citations":[{"why":"This is the original Living Review article, the resource this paper freezes and replaces.","marker":"[1]"},{"why":"This is the online Living Review itself, whose citation patterns illustrate the credit and versioning problems the new model repairs.","marker":"[2]"},{"why":"This is the archival snapshot of the frozen review, the concrete result of the transition and a citable final state.","marker":"[3]"},{"why":"This is a neighboring community-curated resource for machine learning in cosmology, cited as evidence that curation is replacing exhaustive bibliography.","marker":"[4]"},{"why":"This is a curated resource for astronomical data science, further evidence of the same shift.","marker":"[5]"},{"why":"This is a curated resource for HEP software, showing existing community entry points the Guide can link to instead of duplicating.","marker":"[6]"},{"why":"This is a dedicated community resource for simulation-based inference, illustrating the topic-specific ecosystem that makes comprehensive lists less necessary for orientation.","marker":"[8]"},{"why":"This is a community guide to reusable ML models in LHC analyses, serving as a concrete model for the Living Guide's structure.","marker":"[13]"},{"why":"This is a topic-specific review of unfolding with machine learning, cited as evidence that subfields have grown their own reference literature.","marker":"[20]"}],"fun_headline_variants":["HEP-ML freezes its list, launches a curated field guide","Exhaustive review ends; curated guide begins for HEP-ML","HEP-ML list dies, long live the curated guide","From exhaustive to essential: HEP-ML guide replaces list"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The Guide's design rests on the untested premise, based only on informal conversations with researchers, that newcomers chiefly want a few recommended papers, a handful of tutorials, and a map of subfield connections rather than a comprehensive list.","fun_headline_variants_meta":{"raw":{"variants":["HEP-ML freezes its list, launches a curated field guide","Exhaustive review ends; curated guide begins for HEP-ML","HEP-ML list dies, long live the curated guide","From exhaustive to essential: HEP-ML guide replaces list"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000567,"raw_usage":{"total_tokens":2665,"prompt_tokens":904,"completion_tokens":1761,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":1686}},"tokens_in":520,"tokens_out":1761,"duration_ms":13434,"temperature":1.0,"reasoning_tokens":1686,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:31:58.784274+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Count citations and contributions to the archived Living Review versus the Living Guide over the next two years: if the frozen review continues to be cited at its pre-freeze rate while the Guide attracts few contributors or readers, the claim that near-complete bibliographies are no longer particularly useful would be falsified.","supporting_citations":[{"cited_title":"Feickert, B","cited_arxiv_id":null,"evidence_quote":"This is the online Living Review itself, whose citation patterns illustrate the credit and versioning problems the new model repairs."},{"cited_title":"Krause, R","cited_arxiv_id":null,"evidence_quote":"This is the archival snapshot of the frozen review, the concrete result of the transition and a citable final state."},{"cited_title":"Stein,Machine Learning in Cosmology, https://github.com/georgestein/ml-in-cosmology","cited_arxiv_id":null,"evidence_quote":"This is a neighboring community-curated resource for machine learning in cosmology, cited as evidence that curation is replacing exhaustive bibliography."},{"cited_title":"Gully-Santiago,Awesome Astrodata,https://github.com/gully/awesome-astrodata","cited_arxiv_id":null,"evidence_quote":"This is a curated resource for astronomical data science, further evidence of the same shift."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This is a curated resource for HEP software, showing existing community entry points the Guide can link to instead of duplicating."},{"cited_title":"Cranmer and J","cited_arxiv_id":null,"evidence_quote":"This is a dedicated community resource for simulation-based inference, illustrating the topic-specific ecosystem that makes comprehensive lists less necessary for orientation."}],"review_version":1}