{"id":"7621bb71-29c9-487c-bf16-e738dae4404d","arxiv_id":"2605.27404","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"Analysis of 147,074 publications shows AI-assisted writing correlates with smaller, younger research teams that maintain or increase scientific impact.","lead":"This paper finds that research teams using AI-assisted writing tend to be smaller and younger, yet still produce highly cited work. A smart generalist might read it to understand how AI tools are restructuring scientific collaboration and what that means for research policy.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The AI usage detector is the single load-bearing measurement; if it systematically flags certain human writing styles (e.g., formulaic or non-native English prose) as AI-modified, and those styles correlate with author career stage or team composition, all three headline findings could be artifacts.","rationale":"The reader correctly identified the AI detector's unvalidated accuracy as the weakest assumption. I agree this is the single most load-bearing concern: every downstream association—team age, team size, impact—depends on the detector correctly distinguishing AI-modified text from human text. If the detector has systematic false positives correlated with author demographics or writing styles, all three headline findings could be spurious. The reader also identified several secondary concerns (causal framing, Nature PSM non-significance for team size, career age measurement limited to PLoS/Nature, short citation windows). These are real but less load-bearing than the detector validity issue. The career age measurement limitation (measuring seniority only within PLoS/Nature) is a notable secondary concern that could independently bias the 'younger teams' finding, but it affects only one of three claims whereas the detector issue affects all three. The verdict should remain CONDITIONAL: the findings are interesting and the methodology is reasonable, but the detector validation gap prevents full acceptance. The concrete test I propose—examining pre-GPT false positives for systematic demographic patterns—is feasible with existing data and would directly settle whether the concern lands.","tokens_in":23188,"tokens_out":3364,"duration_ms":84885,"concrete_test":"Examine the ~5% of pre-GPT papers that score above the 95th-percentile threshold (these are known false positives). Test whether their author characteristics—career age, team size, discipline, journal—systematically differ from pre-GPT papers scoring below threshold. If, for example, pre-GPT false positives have significantly younger teams or cluster in specific disciplines, the detector has demographic/style bias and the headline associations are likely artifacts of writing style rather than AI adoption. Additionally, manually inspect 50–100 of these false-positive papers to identify what textual features trigger the detector.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claims rest entirely on the AI usage score from Liang et al. (2025). Papers are classified as 'AI-assisted' if their score exceeds the 95th percentile of the pre-GPT distribution for the same journal (§3.2.1). The paper acknowledges that 5% of pre-GPT papers exceed this threshold—these are pure false positives, since ChatGPT did not exist when they were written. The critical question is what makes those pre-GPT papers score high. If the detector flags certain human writing styles—formulaic academic prose, non-native English patterns, or discipline-specific conventions—then post-GPT papers scoring above threshold may be flagged because of how their authors write, not because they used AI. This matters because: (1) junior researchers and non-native English speakers may produce prose that more closely resembles LLM output (simpler syntax, more predictable word choices), creating a spurious association between 'AI usage' and younger teams; (2) if certain disciplines or team compositions correlate with these writing styles, the 'smaller teams' finding could also be an artifact; (3) the 'higher impact' finding could arise if better-organized, more carefully formatted manuscripts (which may trigger the detector) also receive more citations. The paper provides no independent validation of the detector on the PLoS or Nature corpora, nor does it examine whether the pre-GPT false positives share systematic characteristics. The PSM and regression controls address observed confounders but cannot correct for measurement error in the treatment variable itself.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper examines whether AI-assisted writing is associated with changes in research team structure (age and size) and scientific impact, using 147,074 publications from the PLoS family and Nature portfolio (2020–2025). AI usage is measured via the Liang et al. (2025) text-based detection algorithm applied to full-text articles. The authors find that teams with higher AI usage scores tend to be younger and smaller, and that AI-assisted publications are more likely to achieve top-5% FWCI. The analysis employs OLS, quantile regression, Poisson regression, logistic regression, and propensity score matching, with extensive controls and fixed effects.","tokens_in":23325,"tokens_out":2348,"duration_ms":70500,"significance":"The paper addresses a timely and important question at the intersection of AI adoption and scientific team dynamics. Its strengths include the use of full-text (rather than abstract-only) analysis across two distinct publisher portfolios, a multi-method analytical strategy with consistent specifications, and a falsifiable set of predictions that are tested across multiple estimators. The finding that AI-assisted teams are smaller and younger without apparent loss of impact, if robust, has genuine policy relevance. The use of an externally published detector (Liang et al. 2025, Nature Human Behaviour) rather than a home-grown measure is a reasonable methodological choice, though it introduces dependencies discussed below.","major_comments":[{"comment":"§3.2.1: The AI usage score from Liang et al. (2025) is the sole basis for the key independent variable, and the paper acknowledges that 5% of pre-GPT papers exceed the 95th-percentile threshold—these are necessarily false positives since ChatGPT did not exist. The paper does not examine what characteristics these pre-GPT false-positive papers share. If the detector systematically flags certain human writing styles (e.g., formulaic prose, non-native English patterns) that correlate with author career stage or team composition, the headline findings could be partially or wholly artifactual. The authors should at minimum: (a) analyze the pre-GPT false-positive papers for systematic correlates (discipline, author seniority, team size, journal), and (b) discuss the detector's known sensitivity to writing-style confounds. The discipline and journal fixed effects in the regressions partially,但不","section":null},{"comment":"§3.2.2, Eq. (1): Career age is defined as years since an author's first publication found in the PLoS family and Nature Portfolio specifically. A researcher who published extensively in Science, Cell, Lancet, or other venues before appearing in PLoS/Nature would have their career age systematically underestimated. This measurement choice is load-bearing for RQ1 (team age) and could bias results if, for example, AI-assisted writing is more common among researchers who are newer to PLoS/Nature specifically but not necessarily junior in their careers. The authors should either justify why this measurement is adequate given that OpenAlex (which they already use) provides broader publication histories, or acknowledge this as a substantive limitation rather than a minor one.","section":null},{"comment":"Table 2, Panel B (Nature): The PSM result for team size in the Nature portfolio is not statistically significant (ATT = -0.131, p = 0.384). Yet the headline claim in the abstract and §4.2 states that 'teams using AI-assisted writing tend to be smaller' as a general finding across both datasets. The Poisson regression (Table 3) does show a significant coefficient for Nature, but the PSM—the authors' preferred quasi-causal method—does not corroborate this for Nature. The paper should more carefully distinguish which findings are robust across methods and datasets and which are not, particularly in the abstract and conclusion.","section":null}],"minor_comments":[{"comment":"§3.2.3 is labeled '3.3.3 Team size' in the text, breaking the section numbering sequence.","section":null},{"comment":"Table 3: The variable 'First Author Co-author Avg. Career Age (ln)' appears in the Poisson regression but is not described in the variable definitions section. It is unclear how this differs from the 'First Author Prior Team Age (ln)' used in Table 1.","section":null},{"comment":"§4.1.1: The text states that for Nature, 'the differences in team age distributions across the three groups are not statistically significant,' yet the OLS regression (Table 1) and PSM (Table 2) show significant results. This apparent tension between the descriptive and inferential results should be clarified.","section":null},{"comment":"Abstract: The phrase 'higher probability of research teams that employed AI-assisted writing producing highly impactful publications' is grammatically awkward; consider revision.","section":null},{"comment":"The paper would benefit from reporting the actual AI usage score distribution statistics (mean, median, SD) for each of the three groups, beyond the visual distribution in Fig. 1.","section":null},{"comment":"Several references have 2026 publication dates (e.g., Hao et al., 2026; He & Bu, 2026). If these are forthcoming or in-press, this should be noted.","section":null}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the AI detector is legitimate and is the primary reason for the major revision recommendation. The detector is externally validated (published in Nature Human Behaviour), which provides some defense, but the authors' application to a specific corpus without any local validation or analysis of false-positive characteristics is a gap that should be addressable within one revision. The career age measurement issue is also important but may be harder to fully resolve; at minimum the authors should discuss it as a substantive limitation. The paper is otherwise thorough and well-structured, and I would expect it to be publishable after revision."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. All three major comments identify legitimate issues. We will (1) analyze pre-GPT false-positive papers for systematic correlates and add a discussion of detector writing-style confounds, (2) recompute career age using OpenAlex's broader publication histories and treat the current measurement as a robustness check, and (3) revise the abstract, §4.2, and conclusion to accurately reflect that the team-size finding is robust across methods for PLoS but that PSM does not corroborate the Poisson result for Nature. We view all three revisions as strengthening the paper.","responses":[{"response":"The referee raises a valid and important concern. We agree that if the detector systematically flags particular human writing styles that correlate with team composition, our findings could be confounded. We will address this in two ways. First, we will conduct the requested analysis of pre-GPT false-positive papers: we will compare the 5% of pre-GPT papers exceeding the threshold to other pre-GPT papers on discipline, author seniority (career age), team size, and journal. If the false positives are systematically associated with younger or smaller teams, this would suggest a writing-style confound that inflates our headline findings; if not, the concern is substantially mitigated. Second, we will add a dedicated paragraph in §3.2.1 discussing the detector's known sensitivity to formulaic prose and non-native English writing patterns, drawing on the validation evidence reported in Liang et al. (2025) and related literature. We will also note that our regression specifications include discipline and journal fixed effects, which partially absorb systematic differences in writing conventions across fields and venues, though we agree these do not fully eliminate individual-level confounds within a discipline-journal cell. We acknowledge that this is a genuine limitation of relying on any text-based detector and will state this explicitly.","revision_made":"yes","referee_comment":"§3.2.1: The AI usage score from Liang et al. (2025) is the sole basis for the key independent variable, and the paper acknowledges that 5% of pre-GPT papers exceed the 95th-percentile threshold—these are necessarily false positives since ChatGPT did not exist. The paper does not examine what characteristics these pre-GPT false-positive papers share. If the detector systematically flags certain human writing styles (e.g., formulaic prose, non-native English patterns) that correlate with author career stage or team composition, the headline findings could be partially or wholly artifactual. The authors should at minimum: (a) analyze the pre-GPT false-positive papers for systematic correlates (discipline, author seniority, team size, journal), and (b) discuss the detector's known sensitivity to writing-style confounds."},{"response":"The referee is correct that restricting career age to first appearance in PLoS/Nature systematically underestimates seniority for researchers whose first publications appeared in other venues. This is a substantive measurement concern, not a minor one. We already use OpenAlex for author disambiguation and metadata, and OpenAlex does index broader publication histories across venues. We will recompute career age using each author's first publication year as recorded in OpenAlex across all indexed venues, which provides a far more accurate measure of academic seniority. We will then re-estimate all team-age analyses (OLS, quantile regression, PSM) with the revised measure. We will retain the original PLoS/Nature-based measure as a robustness check and report both sets of results. If the pattern holds under the broader measure, this strengthens our findings; if it weakens, we will report that honestly. We will also move this measurement issue from its current implicit treatment to an explicit discussion in the limitations section, acknowledging that even OpenAlex coverage is not complete, particularly for researchers in non-English-dominant scientific communities.","revision_made":"yes","referee_comment":"§3.2.2, Eq. (1): Career age is defined as years since an author's first publication found in the PLoS family and Nature Portfolio specifically. A researcher who published extensively in Science, Cell, Lancet, or other venues before appearing in PLoS/Nature would have their career age systematically underestimated. This measurement choice is load-bearing for RQ1 (team age) and could bias results if, for example, AI-assisted writing is more common among researchers who are newer to PLoS/Nature specifically but not necessarily junior in their careers. The authors should either justify why this measurement is adequate given that OpenAlex (which they already use) provides broader publication histories, or acknowledge this as a substantive limitation rather than a minor one."},{"response":"The referee is correct. The PSM result for team size in the Nature portfolio is not statistically significant, and our current framing in the abstract and §4.2 overstates the generality of the team-size finding. We will revise the manuscript as follows. First, in §4.2, we will explicitly state that the team-size finding is robust across PSM and Poisson regression for PLoS, but that for Nature, the Poisson coefficient is significant while the PSM ATT is not (p = 0.384). We will offer a substantive interpretation: the Nature sample has far fewer AI-assisted papers (1,581 vs. 11,914 in PLoS), reducing PSM statistical power, and Nature teams are on average larger (mean ~8.5 vs. ~6 in PLoS), meaning the absolute difference of 0.131 authors may be too small to detect reliably. Second, we will revise the abstract to state that AI-assisted writing is associated with smaller teams in PLoS across multiple methods, and with a significant negative coefficient in Poisson regression for Nature, while noting that the PSM result for Nature does not reach significance. Third, the conclusion will be adjusted to reflect this nuance rather than presenting the team-size finding as uniformly robust across both datasets. We agree that precision here is essential for the paper's credibility.","revision_made":"yes","referee_comment":"Table 2, Panel B (Nature): The PSM result for team size in the Nature portfolio is not statistically significant (ATT = -0.131, p = 0.384). Yet the headline claim in the abstract and §4.2 states that 'teams using AI-assisted writing tend to be smaller' as a general finding across both datasets. The Poisson regression (Table 3) does show a significant coefficient for Nature, but the PSM—the authors' preferred quasi-causal method—does not corroborate this for Nature. The paper should more carefully distinguish which findings are robust across methods and which are not, particularly in the abstract and conclusion."}],"tokens_in":23119,"tokens_out":1448,"duration_ms":86254,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The headline finding is that papers flagged as AI-assisted tend to come from younger and smaller teams, and these papers are also more likely to be highly cited. The dataset is large (147K publications across PLoS and Nature), the methods are thorough (OLS, Poisson, logistic, quantile regression, PSM with exact discipline matching), and the authors are transparent about the observational design. This is a genuine new contribution at the team level — prior work on AI in science has focused on individual productivity or macro-level trends, not on how team composition correlates with AI usage. Credit is earned for the multi-method approach and for running the analysis on two independent corpora with consistent results on team age and impact. The PLoS-Nature comparison is useful: the team-age effect shows up in both, while the team-size effect is significant in PLoS but not in Nature PSM (p=0.384), which the authors report honestly. The stress-test concern about the AI detector is the right thing to worry about, and it lands. The Liang et al. detector is the single load-bearing measurement. If it systematically flags formulaic or non-native English prose as AI-modified, and those styles correlate with junior authors or smaller teams, the headline associations could be partly or wholly artifactual. The paper does not validate the detector on its own corpora, nor does it examine what characteristics the pre-GPT false positives share. This is the central weakness, and it is not minor — it is the measurement on which everything else depends. That said, the authors did not build the detector themselves, it comes from a published Nature Human Behaviour paper, and the thresholding strategy (95th percentile of pre-GPT distribution per journal) is reasonable as far as it goes. The career-age measurement is a secondary concern: it is computed only within PLoS and Nature publications, so authors with long histories in other venues will appear artificially junior. This biases toward finding younger teams among AI-assisted papers, though the direction and magnitude of the bias is hard to gauge without external validation. The causal framing in the prose occasionally overreaches — phrases like 'AI is pushing back against a long-standing trend' imply more than the observational design supports. The authors acknowledge this limitation in their section on limitations, but the framing in the introduction and discussion could be tighter. This paper is for scientometrics researchers and science-policy scholars interested in how AI is changing team structures. It raises a question worth pursuing seriously, even if the current evidence is correlational and measurement-sensitive. It deserves a serious referee who can push hard on the detector-validation question and the career-age measurement. I would accept it for peer review.","headline":"Large-scale correlational study linking AI-assisted writing to younger, smaller, higher-impact research teams. The association is real and consistently observed, but the measurement rests entirely on an external AI detector whose systematic biases on this specific corpus are unexamined.","tokens_in":24180,"tokens_out":643,"would_cite":false,"duration_ms":53748,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"AI-Assisted Writing Shrinks Research Teams and Raises Their Impact","keywords":[],"falsifier":"If an independent, validated AI-detection method applied to the same corpus produced substantially different group classifications — for instance, if many papers currently labeled 'AI-assisted' were reclassified as human-written — the observed associations with team age, size, and impact could weaken or disappear.","tokens_in":23219,"feed_emoji":"🤖","tokens_out":1203,"duration_ms":32293,"temperature":0.7,"pith_summary":"This paper argues that the adoption of AI-assisted writing in scientific research is reshaping the structure of research teams: teams that use AI tend to be smaller and younger, yet they produce research with higher scientific impact, not lower. Drawing on 147,074 full-text publications from the PLoS family and the Nature portfolio published since 2020, the authors classify papers into AI-assisted and human-written groups using a text-based detection algorithm that estimates the proportion of AI-modified content in each paper. They then compare team age (average career seniority of authors), team size (number of co-authors), and scientific impact (field-weighted citation impact, or FWCI) across these groups using OLS, Poisson, logistic, and quantile regression, plus propensity score matching. The central finding is a triple association: higher AI usage scores correlate with younger teams, smaller teams, and a higher probability of landing in the top 5% of citation impact. The authors interpret this as evidence that AI is partially reversing the decades-long trend toward ever-larger research teams in science, by absorbing writing and editing labor that previously required senior co-authors, while simultaneously freeing researchers to concentrate on the intellectual substance of their work.","feed_headline":"AI-Assisted Writing Shrinks Research Teams and Raises Their Impact","feed_subtitle":"Analysis of 147,000 publications shows AI-using teams are younger and smaller yet more likely to produce highly cited work — a reversal of a","key_machinery":"The paper relies on an AI Usage Score (0–1) derived from a text-based detection algorithm (Liang et al., 2025) that estimates the proportion of AI-modified content in each paper's full text. Papers scoring above the 95th percentile of the pre-ChatGPT distribution for their journal are classified as 'AI-assisted.' The dependent variables are Team Age (average career age of co-authors, based on years since first publication), Team Size (author count), and Top 5% FWCI (a binary indicator for whether a paper falls in the top 5% of field-weighted citation impact). Controls include first-author and corresponding-author prior productivity, collaborator counts, prior team-size and team-age habits,FW","core_discovery":"The paper's central claim is that AI-assisted writing is associated with a structural shift in research teams — toward fewer authors and lower average seniority — and that this shift is accompanied by an increase, not a decrease, in the probability of producing highly cited work. In the PLoS dataset, AI-assisted teams had a 3.0 percentage-point higher probability of producing a top-5% FWCI publication than matched controls; in the Nature dataset, the advantage was 2.7 percentage points. The team-age effect was strongest among the youngest teams (25th percentile), suggesting AI acts as an equalizer for junior researchers who lack the writing experience and academic-discourse mastery of their,","pith_inferences":[],"forward_implications":["If AI-assisted writing genuinely enables smaller, younger teams to produce high-impact research, funding agencies may need to reconsider evaluation criteria that implicitly reward large, senior-heavy teams.","The reversal of the decades-long trend toward larger teams could change the structure of scientific labor markets, reducing demand for senior co-authors whose primary contribution is manuscript polishing and editing.","Junior researchers and non-native English speakers may gain disproportionate advantage from AI writing tools, potentially democratizing access to elite publication venues.","If the pattern generalizes beyond PLoS and Nature, it could signal a broad reorganization of how scientific collaboration is structured, with AI absorbing coordination and communication overhead that previously required larger teams."],"fun_headline_variants":["AI Writing Tools Linked to Smaller, Younger Research Teams With Higher Impact","Smaller Teams, Bigger Impact: How AI-Assisted Writing Reshapes Research","AI-Assisted Writing Tied to Compact, Junior Teams and Highly Cited Papers","Junior Researchers Gain Edge as AI Writing Tools Reshape Team Dynamics"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The entire analysis rests on the accuracy of a single AI-detection algorithm (from Liang et al., 2025) that estimates how much of each paper's text was modified by AI. If this detector systematically misclassifies certain writing styles, disciplines, or author demographics as 'AI-assisted' or 'human-written,' every downstream association — team age, team size, impact — could be biased. The paper does not independently validate the detector on its specific PLoS and Nature full","fun_headline_variants_meta":{"raw":{"variants":["AI Writing Tools Linked to Smaller, Younger Research Teams With Higher Impact","Smaller Teams, Bigger Impact: How AI-Assisted Writing Reshapes Research","AI-Assisted Writing Tied to Compact, Junior Teams and Highly Cited Papers","Junior Researchers Gain Edge as AI Writing Tools Reshape Team Dynamics"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":666,"prompt_tokens":598,"completion_tokens":68,"prompt_tokens_details":null},"tokens_in":598,"tokens_out":68,"duration_ms":53361,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-04T17:47:34.975215+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If an independent, validated AI-detection method applied to the same corpus produced substantially different group classifications — for instance, if many papers currently labeled 'AI-assisted' were reclassified as human-written — the observed associations with team age, size, and impact could weaken or disappear.","supporting_citations":[],"review_version":1}