{"id":"497330c5-0e68-4ab6-9ed4-2cfd0170c2c2","arxiv_id":"2505.12039","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The paper defines a five-level AI4SoS automation hierarchy and demonstrates a preliminary LLM multi-agent society that partially reproduces known correlations between team diversity and citation impact.","lead":"This paper argues that AI can become the foundation of Science of Science research by automating pattern discovery and simulating entire research societies. It also describes a preliminary million-agent LLM simulation that partially reproduces known citation-diversity correlations, but leaves the core claims as a research agenda rather than a validated result.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulated patterns may be inherited from 2010-2020 agent features that overlap the 2010-2011 validation window; without a holdout or permutation control, the proof-of-concept does not show emergent AI-driven replication.","rationale":"The reader's identified weakest assumption is exactly the load-bearing concern: the proof-of-concept's only empirical support is the claimed replication of real-world citation correlations, and that replication may be an artifact of author features drawn from 2010-2020, overlapping the 2010-2011 validation period. My analysis agrees, with the sharper observation that Table 4's co-author lists feed the actual validation-era collaboration network into agent initialization, making the simulated team compositions and citation outcomes potentially predetermined. This does not undermine the paper's forward-looking perspective, the five-level autonomy hierarchy, or the honest limitations section; those remain valuable. But the abstract's 'showcasing' claim rests on the empirical demonstration, and that demonstration lacks the leakage controls needed to distinguish emergent behavior from input-feature recapitulation. The fix is inexpensive: a temporal holdout or permutation control, plus code and data release. Since the reader already conditioned acceptance on these additions, the verdict remains CONDITIONAL rather than being moved to accept or reject.","tokens_in":20362,"tokens_out":9622,"duration_ms":104802,"concrete_test":"Re-run the simulation with author features (ethnicity, affiliation, initial citation count, co-author list) built only from papers published 2002-2009, keeping the pipeline otherwise identical, and compare the three correlations in Figs. 6-7. If the simulated patterns disappear or lose significance, the original result is due to leakage from the 2010-2020 feature window; if they persist, agent behavior is the driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The supporting claim in Sec. 5.3 is that the multi-agent system can 'replicate and uncover key patterns in scientific research,' evidenced by simulated correlations matching real 2010/2011 data (Figs. 6-7). The load-bearing condition is that these correlations emerge from agent behavior, not from the input features. That condition is not secured. Table 4 initializes every agent's citation count, co-author list, ethnicity, and affiliation from OAG papers published between 2010 and 2020, which includes the 2010-2011 validation window. The co-author lists are especially problematic: agents enter the simulation with the real future collaboration network, so the 'Collaborator Selection' step can reassemble the very teams whose 2010/2011 citation patterns are the validation target. Simulated paper citation counts may then reflect the pre-assigned citation impact of those co-authors, and the ethnicity/affiliation correlations are inherited from the input joint distribution rather than produced by the simulated research process. The paper provides no leakage-control experiment, no null baseline, and no sensitivity analysis; one of the three target correlations (affiliation diversity) is already non-significant. Without a test that decouples input feature correlations from simulated outcomes, the abstract's claim that the system 'showcases AI's ability to replicate real-world research patterns' is not supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a forward-looking perspective on using AI to automate Science of Science (SoS) research. It defines AI4SoS, distinguishes it from AI for Science, proposes a five-level autonomy hierarchy (Level 0 through Level 4), surveys open problems in forecasting research trends and understanding research-society dynamics, and discusses challenges such as data bias, system construction, evaluation, and explainability. The paper's empirical core is a proof-of-concept in Sec. 5: a large-scale multi-agent system built on OAG data and LLaMA3.1-8B that simulates one million scientist-agents over 40 epochs. The authors claim that the simulation replicates real-world correlations between citation counts and ethnicity diversity, affiliation diversity, and affiliation ranking, thereby demonstrating that AI can automate pattern discovery and provide a sandbox for SoS experiments. Sections 6-8 provide alternative views, outlook, and conclusions.","tokens_in":20536,"tokens_out":2834,"duration_ms":30701,"significance":"If the central claim were fully supported, the manuscript would make a useful contribution by offering a conceptual framework for AI4SoS and by showing that LLM-based agent simulations can scale to a million agents while producing plausible research-society dynamics. The five-level autonomy hierarchy and the discussion of evaluation and causality are reasonable organizing contributions, and the engineering achievement of a million-agent asynchronous LLM simulation is nontrivial. However, the proof-of-concept's validation is the load-bearing part of the claim that the system can 'replicate and uncover key patterns in scientific research,' and that validation is currently inadequate because of feature-vs-validation-window leakage and the absence of control experiments. The paper also responsibly acknowledges limitations such as missing career trajectories and funding/policy influences, and it includes a peer-review prompt in Appendix B, but these do not compensate for the missing leakage controls. No code or data is provided, so the empirical results are not independently checkable.","major_comments":[{"comment":"The claimed replication is vulnerable to target leakage. Table 4 states that each agent's citation count, co-author list, ethnicity, affiliation, affiliation ranking, discipline, and research topic are extracted from OAG papers published between 2010 and 2020, while the validation patterns are measured on real papers from 2010 and 2011 (Sec. 5.1 and Sec. 5.2). Because the agent features include the validation period and even future collaborations, the simulated correlations between citation counts and ethnicity/affiliation diversity could be inherited from the input joint distribution rather than generated by agent behavior. The co-author lists are especially problematic: the 'Collaborator Selection' stage can reassemble teams from the real 2010-2020 collaboration network, and the simulated papers' citation counts may then reflect the pre-assigned citation impact of those co-authors. The paper provides no leakage-control experiment, no holdout-based feature extraction, and no permutation or null baseline to rule out this inheritance. Without such a control, the abstract's statement that the system 'showcases AI's ability to replicate real-world research patterns' is not supported.","section":"§5.1, Table 4; §5.2; §5.3"},{"comment":"The statistical basis for the replication claim is too weak. The paper reports only qualitative agreement between real and simulated scatter plots, plus a single p-value for affiliation diversity, and that p-value is greater than 0.05. There are no confidence intervals, no R-squared values, no correlation coefficients with uncertainty, and no statement of how many simulated papers the scatter plots are based on. The caption of Fig. 6 refers to 'Strong correlations' in real data, but the text says the simulated correlations are 'slightly weaker,' and no quantitative comparison is given. The authors should report effect sizes and uncertainty for all three relationships, and should state the number of papers and the exact statistical test used for each comparison.","section":"§5.3, Figs. 6-7"},{"comment":"The simulation's ability to 'uncover key patterns' is not distinguished from the reproduction of correlations already present in the input features. Ethnicity diversity, affiliation diversity, and average university ranking are attributes of the agents rather than outcomes of the simulated research process, and the team-size distribution is fit to OAG data (Fig. 5). A minimal control would be to randomize the ethnicity and affiliation labels, or to permute the citation-count assignments, and then verify that the simulated correlations disappear. The authors should also report a sensitivity analysis over the free parameters (team-size distribution parameters, peer review threshold, agent interaction hyperparameters, simulation scale and duration) to show that the observed correlations are not artifacts of specific parameter choices. Without such controls, the claim that the system 'replicate[s] and uncover[s] key patterns' conflates input-feature correlations with emergent AI-driven discovery.","section":"§5.1, Table 4; §5.2, Fig. 5; §5.3"}],"minor_comments":[{"comment":"There is a typo in the phrase 'W e take the position' near the end of the introduction; it should read 'We take the position.'","section":"Sec. 1"},{"comment":"The table says the year of initial-database papers is set to -1, but the text in Sec. 5.1 says the reference database contains papers from 2002 to 2009. Please clarify what -1 means in the simulation timeline and how it interacts with the epoch-based calendar.","section":"§5.1, Table 5"},{"comment":"The x-axis labels '1060 320 1000' are unclear; they appear to be comma-separated numbers rendered without separators. Please label the axis clearly and state the units for all three values.","section":"Fig. 4"},{"comment":"The text says 'the pattern observed in the simulation is not statistically significant' but does not state which test was used or which of the three correlations it refers to. Please clarify the test, the sample size, and the direction of the non-significant result.","section":"§5.3"},{"comment":"No code, configuration files, or simulation output data are provided, which makes the empirical results difficult to assess or reproduce. I encourage the authors to release the agent-initialization pipeline and the analysis scripts.","section":"General"},{"comment":"The caption says 'Strong correlations observed in real data are partially reproduced,' but the text in Sec. 5.3 says the simulated correlations are 'slightly weaker.' Please make the caption consistent with the reported effect sizes.","section":"Fig. 6 caption"}],"recommendation":"major_revision","confidential_remarks":"This is primarily a perspective paper with a proof-of-concept empirical section. The referee report focuses on the proof-of-concept because the abstract and Sec. 5.3 make a strong claim about AI's ability to replicate real-world research patterns. The leaked feature-vs-validation-window overlap is a standard correctness issue that the authors can fix with a holdout design and permutation controls. Given the paper's positioning as a vision/position paper, I would not reject it on the strength of the empirical deficiencies alone, but the overclaim in the abstract must be tempered or the experiments must be redone. I also note that several references are duplicated (e.g., refs. [42] and [100], [43] and [101]), which is a minor editing issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time as a perspective piece, but read the demo with the leakage question in the front of your mind.\n\nWhat is actually new: the five-level AI4SoS autonomy hierarchy is a routine taxonomy, but it is a useful organizing frame for a fragmented discussion. The proof-of-concept is a genuine system artifact—a million-agent LLM simulation with a review and indexing system, built on OAG author data—and running that at scale is not trivial. The paper also does something refreshingly honest: it includes an Alternative Views section and explicitly names limitations in Sec. 5.3, including the non-significant affiliation-diversity result.\n\nWhere it goes soft: the abstract says the system 'showcases AI's ability to replicate real-world research patterns,' and that is not supported. The stress-test concern lands. Table 4 initializes agent citation counts, co-author lists, ethnicity, and affiliation from OAG papers published 2010–2020, and the validation targets are measured on 2010–2011 papers. The co-author lists are especially damning: agents enter the simulation with the real future collaboration network, so the collaborator-selection step can reassemble the very teams whose 2010/2011 citation patterns are the target. The ethnicity and affiliation correlations may simply be inherited from the feature joint distribution rather than produced by the agents' behavior.\n\nThe paper gives no null baseline, no permutation or holdout control, no sensitivity analysis, and no confidence intervals or effect sizes. Those are addressable fixes, not fatal flaws. But as printed, the demo does not establish emergent pattern discovery.\n\nI agree with the reader's conditional verdict. The agenda itself is sensible, and the authors are not overclaiming in the body as much as in the abstract. If they release code and data, run a feature-shuffle or holdout control, and soften the language to 'preliminary illustration,' this becomes a solid early-stage contribution.\n\nWho it is for: AI-for-Science people thinking about simulation infrastructure, and SoS researchers who want to see what LLM multi-agent systems can do. A serious referee should engage with it, but with the requirement that the empirical claim be either fixed or scaled back. I would not cite it for the correlations; I might cite it as an example of the AI4SoS framing if that conversation develops.","headline":"A reasonable agenda paper whose preliminary demo overclaims: the simulated correlations may be inherited from OAG features that overlap the validation window, so treat the proof-of-concept as illustrative until leakage is controlled.","tokens_in":21187,"tokens_out":1159,"would_cite":false,"duration_ms":14587,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI can automate the study of science itself by testing hypotheses in simulated research societies.","keywords":["Science of Science","AI for Science of Science","multi-agent simulation","large language models","research pattern discovery","scientific collaboration","citation analysis","automation hierarchy"],"falsifier":"A decisive test would be to run the same simulation with author features scrambled—for example, randomly permuting ethnicity labels and affiliation rankings across agents—or with agents that retrieve and cite references by random similarity rather than by their learned judgments. If the ethnicity–citation and ranking–citation correlations persist nearly unchanged, they are an artifact of the input data; if they weaken or vanish, the agents' interactions are what generate the patterns.","tokens_in":20054,"feed_emoji":"🧪","tokens_out":9882,"duration_ms":86176,"temperature":0.7,"pith_summary":"This paper argues that artificial intelligence can transform the field known as the Science of Science—the study of how research communities form, collaborate, and produce knowledge—by automating the whole pipeline of data processing, pattern analysis, simulation, and validation. The authors propose a five-level hierarchy of automation, from fully human-driven statistics to fully autonomous AI societies that design and run their own experiments, and they claim that AI offers a 'sandbox' in which hypotheses about science can be tested before being applied to real-world policy. As a proof of concept, they build a multi-agent system of one million large-language-model scientists that choose collaborators, discuss topics, write abstracts, and undergo peer review. After 40 simulated epochs, the citation counts produced by this synthetic community reproduce two real-world patterns: papers with more ethnically diverse author teams attract more citations, and papers from higher-ranked institutions attract fewer citations per output. The result is a first step, in the authors' view, toward making the study of science an experimentally testable discipline.","feed_headline":"1M-agent simulation reproduces real citation patterns","feed_subtitle":"A simulated society of one million researchers reproduces real citation trends.","key_machinery":"The central mechanism is a large-scale multi-agent system in which each scientist is a language-model-driven agent that can communicate, retrieve papers from a shared reference database, and write in natural language. Agents are initialized from a real academic graph—names, affiliations, inferred ethnicity, citation histories, co-author lists, disciplines, and research topics—and the simulation proceeds through collaborator selection, topic discussion, idea generation, novelty assessment, abstract generation, and peer review, with accepted papers entering the database and updating citation counts. The system runs asynchronously so that a society of one million agents can be simulated in about a week, and team sizes are drawn from an exponential distribution fitted to historical data. This machinery is what lets the authors compare simulated citation patterns against real-world patterns and claim that the system can replicate and uncover research dynamics.","core_discovery":"The central claim is that AI can become the foundation of next-generation Science of Science research, not merely as a tool for crunching bibliometric data but as a way to observe research processes in action and validate hypotheses in a virtual world. The paper defines AI for SoS (AI4SoS) as a meta-level enterprise distinct from using AI to solve domain problems: the object of study is the scientific ecosystem itself. The supporting empirical contribution is a preliminary multi-agent system in which scientist agents, initialized with real author characteristics from a large academic graph, form teams, generate ideas, write abstracts, go through peer review, and update citations in a shared database. The authors report that the simulated citation counts reproduce the real-world positive correlation with ethnic diversity and the negative correlation with affiliation ranking observed in 2010 and 2011 data, while the affiliation-diversity correlation is positive but not statistically significant. They take this as evidence that AI-driven simulations have the potential to replicate known patterns and, ultimately, to uncover new ones.","pith_inferences":["An implication the authors leave implicit is that the same leakage concern applies to their scalability claim: if a million-agent simulation merely replays historical co-authorship structures, the 'emergent' patterns are not emergent in a behavioral sense, so the strongest defense of the sandbox idea will require perturbation experiments that change agent incentives and show aggregate patterns shi","The simulation could be extended to probe causal questions that retrospective SoS cannot answer, such as whether the diversity–citation link is driven by team composition itself or by the institutional contexts where diverse teams form; by blocking one pathway in the simulation while controlling the other, one could generate testable hypotheses for real-world data.","A natural next experiment would be to add explicit funding and career-advancement mechanisms to the agents, since the paper names these as missing; measuring how the ethnicity-diversity correlation responds would reveal how sensitive the reproduced pattern is to resource allocation.","The automation hierarchy could be repurposed as an evaluation instrument for the broader AI-for-science movement, mapping existing autonomous-science systems to levels 0–4 and exposing where the real bottleneck lies—simulation realism, validation metrics, or explainability—rather than treating each system in isolation."],"forward_implications":["Researchers in the Science of Science could run controlled experiments on funding rules, team-size distributions, reviewer thresholds, or policy interventions inside a million-agent society before trying them in the real world.","The five-level automation hierarchy gives the field a shared vocabulary and a roadmap, making it clear which stages of research are already automatable and which still require human oversight.","The successful reproduction of the ethnicity-diversity and affiliation-ranking correlations suggests that LLM-driven agents can capture at least some emergent social dynamics of science, opening the door to discovering patterns that are invisible in retrospective statistics.","The failure of the affiliation-diversity correlation to reach significance in the simulation provides a concrete benchmark: any improved AI4SoS system should be expected to reproduce all three real-world correlations, not just the first two.","If fully realized, automated SoS discovery could make science-foresight tools—trend analysis, collaboration recommendations, policy evaluation—available to individual researchers and smaller institutions, not only to large labs with dedicated data teams."],"supporting_citations":[{"why":"Supplies the foundational definition of Science of Science and its core questions, which the paper's vision of AI4SoS is built upon.","marker":"[1]"},{"why":"Provides the real-world finding that ethnic diversity in scientific collaboration is positively associated with citation impact, the pattern the simulation reproduces.","marker":"[11]"},{"why":"Supplies the multi-agent design for scientific idea generation and the peer-review mechanism that the proof-of-concept system extends.","marker":"[22]"},{"why":"Provides the asynchronous, large-scale agent simulation architecture that makes the million-agent society computationally feasible.","marker":"[23]"},{"why":"Supplies the large-scale academic graph dataset used to initialize agents with authorship, affiliation, citation, and discipline information.","marker":"[47]"},{"why":"Supplies the LLM-based automated scientific discovery workflow that the paper adapts to a social-simulation setting.","marker":"[65]"},{"why":"Supplies the name-ethnicity classifier used to infer author ethnicity from names in the dataset.","marker":"[81]"},{"why":"Supplies the citation-count metric for measuring research impact, which the simulation and real-world validation both use.","marker":"[82]"}],"fun_headline_variants":["AI agents recreate real-world citation patterns","One million simulated researchers match real trends","AI simulation validates diversity-citation links","Multi-agent AI models the science ecosystem","AI-driven simulation reproduces citation dynamics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the simulated correlations arise from the agents' own simulated behavior rather than being inherited from the real author characteristics (such as ethnicity, affiliation rank, and past citation counts) that were pre-loaded into the agents from historical data; if those injected features alone explain the correlations, the proof-of-concept would only be echoing its inputs, not discovering anything.","fun_headline_variants_meta":{"raw":{"variants":["AI agents recreate real-world citation patterns","One million simulated researchers match real trends","AI simulation validates diversity-citation links","Multi-agent AI models the science ecosystem","AI-driven simulation reproduces citation dynamics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1239,"prompt_tokens":895,"completion_tokens":344,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":283}},"tokens_in":511,"tokens_out":344,"duration_ms":3868,"temperature":1.0,"reasoning_tokens":283,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:42:25.804624+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would be to run the same simulation with author features scrambled—for example, randomly permuting ethnicity labels and affiliation rankings across agents—or with agents that retrieve and cite references by random similarity rather than by their learned judgments. If the ethnicity–citation and ranking–citation correlations persist nearly unchanged, they are an artifact of the input data; if they weaken or vanish, the agents' interactions are what generate the patterns.","supporting_citations":[{"cited_title":"IEEE Transactions on Knowledge and Data Engineering35(9), 9225–9239 (2022)","cited_arxiv_id":null,"evidence_quote":"Supplies the large-scale academic graph dataset used to initialize agents with authorship, affiliation, citation, and discipline information."},{"cited_title":"In: Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp","cited_arxiv_id":null,"evidence_quote":"Supplies the name-ethnicity classifier used to infer author ethnicity from names in the dataset."},{"cited_title":"Nature communications 10(1), 5170 (2019)","cited_arxiv_id":null,"evidence_quote":"Supplies the citation-count metric for measuring research impact, which the simulation and real-world validation both use."}],"review_version":1}