{"id":"9083ec18-f05e-49a6-8ce5-4a1f53ba0858","arxiv_id":"2608.10715","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"An analysis of 1.19 million biomedical papers estimates that 89% showed signs of LLM-assisted writing by December 2025, using extrapolated word frequencies.","lead":"By the end of 2025, an estimated 89% of open-access biomedical papers in PubMed Central show excess vocabulary associated with large language models, with Discussion sections affected more than Methods sections. The estimate comes from tracking how the frequency of marker words changed after 2022 and could give journals a monitoring tool for AI-assisted writing.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 89% figure is a lower bound whose tightness is unverified; the validation simulation builds in the key assumption, and the paper omits the q/p_human values that would bound the bias.","rationale":"The paper provides a clear, reproducible method and useful section- and country-level analyses, but the central claim of an unbiased 89% point estimate is not supported: the estimator is a lower bound unless p_LLM = 1 for the selected marker set, and the paper's justification is both mathematically incomplete (q close to 1 implies p_LLM ≥ q, not p_LLM ≈ 1 in the needed sense) and empirically unquantified (no reported q, p_human, or bias bound for the chosen thresholds). The simulation 'validates' the procedure under a parameterization that forces the union p_LLM to be near 1, so it does not test the realistic case where the marker set is imperfect. This is the same load-bearing weakness the reader identified, and the appropriate response is to require the authors to either report the bias bound or reframe the headline as a conditional lower bound. Therefore, the reader's CONDITIONAL verdict should stand unchanged.","tokens_in":13471,"tokens_out":10844,"duration_ms":88757,"concrete_test":"Re-run the public code and, for each section and for December 2025, record the selected threshold T, q(T), p̂_human(T), and the implied maximum upward bias Δ = (1 − q)/(1 − p̂_human). If Δ exceeds 0.03 for the full-paper estimate, the headline and abstract must be revised to present 89% as a lower bound (or a range) rather than a point estimate; if Δ ≤ 0.03, the tightness assumption is quantitatively supported and the point estimate is defensible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The estimator β̂ = max_T β̂_LB with β̂_LB = (q − p̂_human)/(1 − p̂_human) is mathematically a lower bound on the true β, because it replaces the unobserved p_LLM by its upper bound 1. Equality β̂ = β requires p_LLM = 1 for the chosen marker-word set: every LLM-assisted paper must contain at least one marker word. The paper asserts that 'q close to 1 (and hence pLLM close to 1)' justifies tightness, but the data only imply p_LLM ≥ q (since q = (1−β)p_h + β p_LLM and β ≤ 1). The residual bias β − β̂_LB = (q − p_h)(1 − p_LLM) / ((p_LLM − p_h)(1 − p_h)) can be as large as (1 − q)/(1 − p_h) even when p_LLM is, say, 0.95. The simulation 'validation' samples p_LLM = (1+δ)p_human with δ ∈ [0.5, 5], which guarantees that the union p_LLM for a broad threshold set is near 1; it therefore validates unbiasedness only under the exact condition at issue. The paper never reports q and p_human for the selected thresholds, so the potential gap between the 0.89 estimate and the true β is unbounded in the presented results: for example, q = 0.98 and p_human = 0.80 gives β̂_LB = 0.90 while the true β could be 1.00. The abstract's '89% of papers show excess' should be read as 'at least 89%' unless the tightness is quantified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a frequency-based estimator for the prevalence of LLM-assisted writing in a corpus. The authors apply it to open-access biomedical papers from PubMed Central, estimating β̂=0.89 for full papers in December 2025, with Discussion sections at 0.68 and Methods at 0.32 when length is controlled via 255-word crops. The method builds on a 379-marker-word list from the authors' prior work, models the counterfactual human usage p̂_human by linear extrapolation of 2018–2022 frequencies, and uses β̂ = max_T (q−p̂_human)/(1−p̂_human) after discarding thresholds with large standard errors. A simulation experiment is reported to support the estimator's accuracy. The paper also compares results across sections and affiliation countries, and frames the result as an unbiased estimate rather than a lower bound.","tokens_in":13844,"tokens_out":5129,"duration_ms":48345,"significance":"If the tightness of the lower bound and the faithfulness of the counterfactual extrapolation were established, this would be a valuable contribution: it would provide the first corpus-level estimate of LLM usage that goes beyond a lower bound, with a transparent public dataset and code (GitHub). The section-level and country-level comparisons are policy-relevant and extend prior work. However, the load-bearing 'unbiased' claim is not currently supported: the estimator remains a lower bound whose gap from the true β is not quantified in real data, and the simulation validates the procedure under the exact assumption (p_LLM close to 1) that is at issue. The core scientific claim is therefore conditional on additional sensitivity analyses and reporting of the unobserved quantities (q and p̂_human for selected thresholds).","major_comments":[{"comment":"The tightness assumption is not logically justified. From q = (1−β)p_h + β p_LLM one can only conclude p_LLM ≥ q, not p_LLM ≈ 1; for example, with q = 0.98 and p_h = 0.80, the estimator returns β̂_LB = 0.90 while the true β could be 1.00, a bias of 0.10 even though q is close to 1. The residual bias is (q−p_h)(1−p_LLM)/((p_LLM−p_h)(1−p_h)), which is not bounded by the reported data. The authors should report q and p̂_human for every selected threshold (Table 1 rows and Figure 1f/g), and provide a sensitivity analysis of β̂ as a function of an assumed p_LLM (e.g., 0.90, 0.95, 0.99) or a formal upper bound on the bias. Until then, the abstract's '89%' should be stated as 'at least 89%'.","section":"§2, Eq. (4) and the sentence 'As this maximum was typically achieved at q close to 1...'"},{"comment":"The simulation draws p_LLM = (1+δ) p_human with δ ~ U(0.5, 5), which by construction drives the union probability over a broad threshold set close to 1, so the simulation validates the estimator only in the regime where p_LLM ≈ 1, i.e., the very condition the paper asserts without real-data support. The simulation does not test the estimator under moderate p_LLM (e.g., 0.8–0.95), where the lower-bound gap is material, nor under a misspecified p_human trend (e.g., quadratic drift or a level shift). Please rerun the simulation with fixed p_LLM values and with a nonlinear counterfactual, and report bias and coverage for β in [0,1]. The current Figure 1h therefore does not substantiate the claim of 'high accuracy' for the real-data scenario.","section":"§4, 'Simulation experiment'"},{"comment":"The paper correctly identifies that β̂ is only valid if p̂_human faithfully estimates p_human in 2025. This is a load-bearing assumption for the 0.89 point estimate, yet no placebo test is provided: for example, fitting the linear extrapolation on 2014–2017 data to predict 2018–2022, or applying the same estimator to a control set of words that are not associated with LLM use, would calibrate the extrapolation error. Without such a demonstration, the excess (q − p̂_human) could reflect non-LLM vocabulary drift or human stylistic adaptation to LLM output, as the authors themselves acknowledge. The authors should either add such a calibration or explicitly present the headline number as a lower bound subject to this assumption.","section":"§3, 'Limitations' paragraph beginning 'Our estimate hinges...'"},{"comment":"Selecting the threshold T that maximizes β̂_LB over a grid introduces an upward finite-sample selection bias: each individual β̂_LB is a valid lower bound under the model, but the maximum of many noisy lower-bound estimates will tend to exceed the true maximum lower-bound curve. The simulation incorporates the same selection and reports negligible bias, but that simulation again relies on near-saturated p_LLM. Please report the number of thresholds within one standard error of the maximum, provide a selection-adjusted standard error (e.g., bootstrap or max-bias correction), or show that the maximum is attained over a plateau rather than at a single noisy point.","section":"§2, 'To find an optimal value of T' and Fig. 1g"}],"minor_comments":[{"comment":"The phrase '89% of papers show excess of LLM-associated vocabulary' should read 'at least 89%' given that the estimator is a lower bound; the same wording appears in the Discussion and should be made consistent.","section":"Abstract"},{"comment":"The notation alternates between 'p human' and 'p_h' (or 'p_human') in Eqs. (2)–(4). Please define one symbol (e.g., p_h for human usage frequency) and use it consistently; also distinguish the estimated p̂_human from the true p_human in the derivation.","section":"§2, displayed equations"},{"comment":"The variance formula for Var[β̂] appears to treat q and p̂_human as independent; the methods text should note when this independence assumption is used and whether the covariance term is negligible in practice.","section":"§4, 'Standard errors'"},{"comment":"The text uses 'Pubmed' in the first line of the Methods section; it should be 'PubMed' for consistency with the rest of the manuscript.","section":"§4, 'PubMed Central'"},{"comment":"Table S1 and S2 contain references without full citation details (e.g., author-year placeholders such as '2024' and '2025' are used in the table cells). Please include the full author list or a footnote mapping the table entries to the reference list.","section":"Supplementary Tables"}],"recommendation":"major_revision","confidential_remarks":"The authors should be aware that the reader's central concern matches the manuscript's own limitation statement: the estimator is a lower bound, and the tightness is asserted rather than demonstrated. The marker-word list inherited from Kobak et al. (2025) with two overlapping co-authors may raise a circularity concern among reviewers; the authors should present an external validation of the marker list or an explicit argument for why the inheritance does not bias the estimate. If the authors can supply the requested q/p̂_human reporting and a sensitivity analysis over p_LLM, the paper would be a credible contribution; without it, the headline 89% figure is not supported as a point estimate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the empirical work is genuinely useful: the section-level and country-level breakdowns, the 255-word crop control, and the full-text PMC analysis are new and will be widely cited. Second, the headline “89%” is not supported as a point estimate. The estimator β̂ = max_T (q − p̂_human)/(1 − p̂_human) is a lower bound, and the paper's claim that it is tight rests on the assumption that the chosen marker-word set has p_LLM close to 1. That assumption is asserted, not demonstrated, and the simulation that claims to validate unbiasedness builds it in by sampling p_LLM = (1+δ)p_human with δ up to 5, which guarantees saturation for large word sets.\n\nThe math itself is clean. The delta-method standard errors look right, the threshold-selection rule is sensible, and the decision to use fixed thresholds for country comparisons is honestly flagged as producing underestimates. The paper also does a good job situating itself against prior lower-bound estimates and survey data; the convergence of the 89% figure with survey self-reports in late 2025 is suggestive.\n\nThe soft spots are real, though. The key one: the gap between β̂_LB and true β depends on how far p_LLM is below 1, and the paper never reports the q and p̂_human values at the selected thresholds. Without those, the reader cannot tell whether the 89% is 89% or, say, 60% with p_LLM = 0.95. The limitation section acknowledges the p̂_human extrapolation risk but does not acknowledge that the tightness assumption is equally load-bearing. The simulation is not evidence of tightness on real data because it generates the very condition needed for unbiasedness.\n\nOne more minor point: the marker-word list is inherited from the authors' prior paper, which is fine, but it means the method is not fully independent of the discovery set.\n\nWho is this for? People working on AI-in-science measurement and research integrity. They will want to read it carefully, and they should cite it with a caveat. It deserves a serious referee: the claim is important, the code is public, and the empirical decomposition is valuable. A good referee will ask for the q/p̂_human diagnostics and a sensitivity analysis on p_LLM, and for the abstract to say “at least 89%” unless the tightness is quantified.\n\nI would send it to review, with the expectation of major revision. The method is promising; the overclaim is correctable.","headline":"Useful empirical breakdowns and a clean lower-bound estimator, but the '89%' point estimate overreaches: tightness is asserted, not shown.","tokens_in":14399,"tokens_out":2035,"would_cite":true,"duration_ms":18721,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By the end of 2025, 89% of open-access biomedical papers carried excess LLM-associated vocabulary, according to a new word-frequency analysis.","keywords":["LLM-assisted writing","biomedical publications","word frequency analysis","marker words","prevalence estimation","PubMed Central","academic integrity","ChatGPT"],"falsifier":"Take a set of biomedical papers published in 2025 whose authors declare no LLM use and whose text passes manual inspection, and compare their marker-word frequency to the linear extrapolation from 2018–2022; if the frequency runs significantly above the trend, the counterfactual is biased and the 89% figure is inflated.","tokens_in":13185,"feed_emoji":"🤖","tokens_out":6111,"duration_ms":53501,"temperature":0.7,"pith_summary":"This paper claims that by December 2025, 89% of open-access biomedical papers in PubMed Central showed signs of LLM-assisted writing or editing, up from about 19% in 2023. The authors argue that earlier frequency-gap methods systematically underestimated the true prevalence, and they offer a way to convert a lower bound into a point estimate by assuming, with simulation support, that the best lower bound is tight. The result matters because it suggests that LLM assistance has become the default in biomedical publishing, with implications for research integrity, peer review, and language equity. The method tracks the share of papers containing any of a fixed set of 'marker words' whose usage rose sharply after ChatGPT's release.","feed_headline":"89% of biomedical papers carry LLM writing markers by 2025","feed_subtitle":"A word-frequency analysis of 1.2 million papers finds LLM editing in nearly all, led by Discussion sections.","key_machinery":"The central identity is the two-component mixture q = (1−β)p_human + β p_LLM, which yields the lower bound β ≥ (q − p_human)/(1 − p_human). The paper's estimate β̂ is the maximum over word-set thresholds of this lower bound, computed with the counterfactual p̂_human from linear extrapolation of 2018–2022 frequencies, justified by simulation. The 379 marker words and the threshold selection for word-set size carry the analysis; the critical step is the 'max lower bound' assumption that the optimal word set has p_LLM close to 1, which converts a bound into a point estimate.","core_discovery":"Using full texts of 1,194,287 open-access biomedical papers from PubMed Central, the authors estimate that by December 2025, 89% of papers contained excess LLM-associated vocabulary. The estimate comes from a set of 379 non-content marker words (e.g., 'these', 'potential', 'delves') identified in prior work. For each candidate word set, they compare the observed share of papers containing any marker word to a counterfactual share extrapolated from the 2018–2022 linear trend. The excess is converted into an LLM-usage prevalence using the mixture identity q = (1−β)p_human + β p_LLM with p_LLM bounded above by 1, and the maximum lower bound over word-set sizes is taken as the estimate. The paper argues this is a tighter and more accurate estimate than earlier frequency-gap or mixture-model approaches, supported by a simulation that recovers the true β.","pith_inferences":["If the marker-word trend continues, the extrapolated counterfactual will become increasingly unreliable, so the 89% figure is a moving target that will need re-estimation with fresh human-written baselines, such as preprints that predate ChatGPT or non-academic writing.","The reported country differences could reflect differences in English proficiency and reliance on LLM translation rather than direct LLM drafting; a follow-up could separate translation-assisted from drafting-assisted writing.","The method could be ported to other corpora, such as grant applications, clinical notes, or policy documents, wherever a pre-LLM baseline exists.","The tightness assumption is testable: analyzing version-control histories of manuscripts would give a direct measure of p_LLM and validate or correct the estimation."],"forward_implications":["If correct, the 89% figure means LLM assistance is now the norm in biomedical publishing, not a minority practice.","Discussion sections are roughly twice as LLM-influenced as Methods (68% vs 32% in length-controlled 255-word crops), suggesting authors use LLMs most for framing and interpretation.","The gap between native-English-majority countries (37%) and others (72%) indicates LLM writing tools are closing language gaps but also creating large cross-country differences in dependence.","Because frequency-based methods estimate direct LLM editing and surveys report similar or higher usage, policies that presume disclosure may need to assume near-universal exposure to LLM-assisted text."],"supporting_citations":[{"why":"Supplies the 379 marker words and the original frequency-gap lower-bound logic that this paper extends into a point estimate.","marker":"Kobak et al., 2025"},{"why":"Provides the earlier frequency-gap estimate on Dimensions and OpenAlex that this paper argues systematically underestimates true LLM usage.","marker":"Gray, 2025"},{"why":"Represents the mixture-model approach that this paper contrasts with, relying on prompt-specific LLM text and yielding variable estimates.","marker":"Liang et al., 2025"},{"why":"Documents human avoidance of LLM marker words, a counter-effect the paper cites in its limitations discussion of the counterfactual assumption.","marker":"Geng and Trotta, 2025"}],"fun_headline_variants":["89% of biomedical papers show LLM writing markers","LLM writing detected in 89% of biomedical papers","9 in 10 biomedical papers carry LLM writing signs","Word-frequency analysis: 89% of biomedical papers have LLM words","By end of 2025, 89% of biomedical papers had LLM writing markers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole estimate rests on believing that the pre-2023 linear trend in how often humans used these marker words would have continued unchanged through 2025 if ChatGPT had never existed, and that among the candidate marker-word sets one achieves near-certain use in LLM-edited text.","fun_headline_variants_meta":{"raw":{"variants":["89% of biomedical papers show LLM writing markers","LLM writing detected in 89% of biomedical papers","9 in 10 biomedical papers carry LLM writing signs","Word-frequency analysis: 89% of biomedical papers have LLM words","By end of 2025, 89% of biomedical papers had LLM writing markers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000901,"raw_usage":{"total_tokens":3871,"prompt_tokens":927,"completion_tokens":2944,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":2854}},"tokens_in":543,"tokens_out":2944,"duration_ms":21327,"temperature":1.0,"reasoning_tokens":2854,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:48:41.539089+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of biomedical papers published in 2025 whose authors declare no LLM use and whose text passes manual inspection, and compare their marker-word frequency to the linear extrapolation from 2018–2022; if the frequency runs significantly above the trend, the counterfactual is biased and the 89% figure is inflated.","supporting_citations":[{"cited_title":"Human-LLM coevo- lution: Evidence from academic writing","cited_arxiv_id":null,"evidence_quote":"Documents human avoidance of LLM marker words, a counter-effect the paper cites in its limitations discussion of the counterfactual assumption."}],"review_version":1}