{"id":"9de33032-5453-40cd-bdfa-dd8e9e6dc2ae","arxiv_id":"2505.06108","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Frontier LLMs now match or beat expert-level performance on several challenging biology benchmarks, while older benchmarks show saturation.","lead":"New tests of 27 AI language models on eight biology exams show the best models now answer harder biology questions as well as or better than human experts. The study also finds that many older biology benchmarks have become too easy to tell models apart.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human expert baselines were measured on different subsets than the model evaluations; on several benchmarks the 'exceeds expert' claim lacks a valid same-set comparison.","rationale":"The reader's weakest assumption correctly identifies the most load-bearing concern: the human baselines are not measured on the same subsets as the model evaluations, and in several cases the sample sizes are so small that the claimed 'exceeds expert' margins are not statistically meaningful. The paper is transparent about this limitation (Section 2.2.2) and even discusses the need for better human baselines (Section 4.1), so the concern is not hidden, but it directly affects the central claim. The VCT-Text result (o3 at 46.1% vs. 22.6% expert baseline) is the headline example, and it depends on the validity of the VCT human baseline; a scoring-rule mismatch or subset mismatch would undermine the 'twice as well' statement. Similarly, CloningScenarios at 61% vs. 60% is within sampling noise, so the 'match or exceed' claim for that benchmark is too strong. The rest of the paper—the systematic 27-model evaluation, the longitudinal trends, the robustness checks across prompting strategies, and the public code—are solid and valuable. The concern does not warrant rejection; it warrants a conditional interpretation, as the reader concluded. The concrete test—same-question, same-scoring comparison with confidence intervals—would settle whether the headline claim holds, and is feasible given the public benchmark data and the author's stated willingness to share raw logs.","tokens_in":32015,"tokens_out":3874,"duration_ms":40625,"concrete_test":"For each benchmark where 'exceeds expert' is claimed, recompute model accuracy on the exact questions used for the human baseline (e.g., GPQA_extended biology questions, RAND's ~20% WMDP-Bio sample, the VCT text questions actually answered by virologists) and compare to the reported expert accuracies with 95% confidence intervals. For CloningScenarios, run a two-proportion significance test (33 questions; 61% vs. 60%); if the confidence interval on the difference includes zero, soften the 'exceeds expert' wording.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that models 'outperform experts' rests on comparing model scores on the exact benchmark subsets used in this study (e.g., GPQA_main Biology, WMDP-Bio, LAB-Bench CloningScenarios, VCT-Text) with human baselines taken from other subsets in the original publications. Section 2.2.2 explicitly concedes that 'baseline measurements were conducted on slightly different subsets than my evaluation.' For GPQA-Bio, the expert baseline of 66.7% is from GPQA_extended biology questions, while models were scored on 78 GPQA_main biology questions; if these sets differ in difficulty, the 81% vs. 66.7% comparison is not apples-to-apples. For WMDP-Bio, the 60.5% baseline comes from two RAND researchers on ~20% of questions, a small, potentially non-representative sample. Most fragile is CloningScenarios: with only 33 questions, the model's 61% vs. the 60% expert baseline is a difference of less than one question, and the expert baseline itself was measured on a small 'representative' sample; no confidence intervals or significance tests are provided. For VCT-Text, the human baseline of 22.6% depends on the scoring rule used; if the VCT human evaluation allowed partial credit while the model evaluation uses exact all-or-nothing matching, the 'twice as well' claim would be confounded. The paper's own Discussion (Section 4.1) calls for better human baselines, acknowledging the problem. Without a same-question, same-scoring comparison with uncertainty, the headline 'models exceed experts' is not rigorously established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic evaluation of 27 large language models released between November 2022 and April 2025 on eight biology-related benchmarks (PubMedQA, MMLU-Bio, GPQA-Bio, WMDP-Bio, LAB-Bench LitQA2, CloningScenarios, ProtocolQA, and VCT-Text), using ten independent zero-shot runs per model-benchmark combination and the Inspect AI framework. The headline empirical findings are large performance gains over time, most strikingly a 4-fold improvement on VCT-Text with o3 at 46.1% versus a 22.6% expert virologist baseline, and claims that several models match or exceed documented expert performance on GPQA-Bio, WMDP-Bio, and CloningScenarios. Secondary analyses compare zero-shot, five-shot, and chain-of-thought prompting on four models and vary reasoning effort for o3-mini and Claude 3.7 Sonnet. The paper concludes that zero-shot evaluation is generally reliable, inference scaling usually helps, several established benchmarks are saturating, and better human baselines are needed.","tokens_in":32292,"tokens_out":5040,"duration_ms":47897,"significance":"The study is a useful independent benchmarking contribution: it applies a consistent protocol across a broad model family, performs repeated runs, reports standard deviations, makes evaluation code publicly available, and connects performance to model release dates and training compute. If the raw scores are taken at face value, the model-to-model time series on GPQA-Bio, VCT-Text, and LAB-Bench is informative for tracking AI capabilities in biology. The central 'outperforms experts' claim, however, is not established at the same standard as the model-to-model comparisons, because it relies on external human baselines measured on different subsets, with different tool access and, in at least one case, a comparison within sampling noise. The Discussion's call for more rigorous human baselines is appropriate but is in tension with the strength of the abstract and title.","major_comments":[{"comment":"Section 2.2.2 states that 'baseline measurements were conducted on slightly different subsets than my evaluation,' and this directly affects the paper's central claim. The GPQA-Bio expert baseline of 66.7% comes from GPQA_extended biology questions, whereas models were evaluated on 78 GPQA_main biology questions; the WMDP-Bio expert baseline comes from two RAND researchers on roughly 20% of WMDP-Bio; and the MMLU-Bio expert baseline of 89.8% is an estimate from exam percentiles across all MMLU categories, not a direct measurement on MMLU-Bio. Comparisons in Figure 2 therefore mix benchmark subsets and estimation methods, so the statement in Section 3.1 that top models 'surpassed expert performance on GPQA-Bio, WMDP-Bio, and CloningScenarios' is not supported at the same level of evidence as the model-to-model results. I would ask the author to either add same-question human baseline measurements for the exact evaluated subsets or to restrict the abstract, title, and Section 3.1 to 'exceeds published expert baselines' with explicit caveats.","section":"2.2.2 and Figure 2"},{"comment":"CloningScenarios consists of 33 questions, and the 'exceed' claim is numerically fragile. In Table A4.1, o3 attains 61.2% ± 4.5% and Claude 3.7 Sonnet (16K) attains 60.9% ± 4.4% against a 60% expert baseline; the difference is less than one question, and no confidence interval is given for the expert baseline, which itself came from a representative sample. Under an exact binomial test, 61% on 33 questions versus 60% is entirely consistent with noise, so Section 3.1's claim of 'slightly exceeding the 60% expert baseline' overstates the evidence. Please report an uncertainty estimate for this comparison and avoid claiming 'exceed' unless it survives a test or is explicitly framed as a point estimate.","section":"3.1 and Table A4.1"},{"comment":"The 'twice as well as expert virologists' claim depends on scoring-rule comparability. Section 2.4 describes using Inspect's choice scorer with multiple_correct=True, which requires an exact all-or-nothing match, but Section 2.2.2 simply reports the VCT human baseline of 22.6% without stating whether the original human evaluation used the same exact-match scoring on the same 101 text-only questions. If the human evaluation allowed partial credit, or if the text-only baseline was computed over the experts' tailored subsets rather than the full VCT-Text set, the 46.1% versus 22.6% comparison would no longer be apples-to-apples. The manuscript should verify and document the original VCT scoring rule and subset, or downgrade the abstract's 'twice as well' claim.","section":"2.4 and 2.2.2 (VCT-Text)"}],"minor_comments":[{"comment":"The table header contains '03-mini' where 'o3-mini' is intended.","section":"Table A4.1"},{"comment":"The sentence 'performance increased 2.6-fold the study period' appears to be missing 'over' before 'the study period.'","section":"Section 3.1"},{"comment":"The explanation of negative normalized values is hard to parse; a concrete example would improve clarity.","section":"Figure 2 caption"},{"comment":"Raw logs being 'available upon reasonable request' is weaker than the reproducibility standard implied by the public code; consider a more specific data-sharing plan.","section":"Section 2.4 and Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The title and abstract make claims stronger than the evidence analyzed in the paper's own Section 2.2.2. The manuscript would be better framed as a longitudinal benchmark report; if the journal values strong headline claims, the baselining issue needs to be resolved first."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuinely useful paper: 27 models, eight biology benchmarks, ten runs each, run through one framework with code public. That consistent longitudinal dataset is the real contribution, and it's a good one. The prompting and reasoning-effort analyses are a nice add, and the author is transparent about model versions and evaluation details. I'd want this data on hand next time I need to compare biology-capability claims across models.\n\nThe soft spot is exactly where the reader put it. The 'exceeds expert' headline rests on human baselines that were not measured on the same subsets as the models. GPQA's expert baseline is from the extended set, models ran on main; WMDP's is two RAND researchers on ~20% of questions; CloningScenarios has 33 questions and the 61% vs 60% gap is literally less than one question. No confidence intervals anywhere on the human side. VCT scoring could differ in how partial credit is handled. The author acknowledges all this in Section 2.2.2 and again in Discussion 4.1, which is to his credit, but the abstract still says 'twice as well as expert virologists' without flagging that the baseline is not matched. That overstates the evidence by a meaningful amount.\n\nThe 'match or exceed' claim is plausible—models clearly are closing the gap—but it is not rigorously established on these numbers. The fix is straightforward: either run a small human evaluation on the exact subsets, or soften the abstract and add uncertainty bounds around the baseline comparisons. The model-to-model and longitudinal findings don't depend on those baselines, and they hold up fine.\n\nThis deserves a serious referee. It is reproducible, the measurements are likely sound, and the benchmark-methodology community will care about the comparison issues. I'd send it out with a request to either match the baselines or temper the central claim. With that revision, it's a solid, citable contribution. My own verdict is skeptical of the headline as written but not of the underlying work.","headline":"Valuable systematic benchmark suite, but the 'outperforms experts' headline outruns the mismatched human baselines.","tokens_in":32813,"tokens_out":1897,"would_cite":true,"duration_ms":20231,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Frontier LLMs match or exceed human biology experts on several challenging benchmarks, with o3 more than doubling the expert virologist baseline on VCT-Text.","keywords":["large language models","biology benchmarks","expert baselines","Virology Capabilities Test","GPQA","benchmark saturation","inference scaling","biosecurity"],"falsifier":"Retest the 101 VCT-Text questions with a panel of PhD virologists under the same no-LLM conditions used in the original baselining, scoring exact multiple-response matches; if their mean accuracy reaches 46.1% or higher, the paper's headline claim that o3 performs twice as well as expert virologists fails on that benchmark.","tokens_in":31782,"feed_emoji":"🧬","tokens_out":9393,"duration_ms":81325,"temperature":0.7,"pith_summary":"This study evaluates 27 large language models released between 2022 and 2025 on eight biology benchmarks, running each model-benchmark pair ten times. It claims that frontier models now match or exceed documented expert-level performance on several of the hardest tests: o3 reaches 46.1% on VCT-Text against a 22.6% expert virologist baseline, and o3 and Claude 3.7 Sonnet reach 61% on CloningScenarios against a 60% expert baseline. The paper also finds that top-model performance grew roughly 2- to 4-fold on these challenging benchmarks over the study period, while PubMedQA, MMLU-Bio, and WMDP-Bio plateaued below 100%, suggesting saturation and noisy answer labels. A sympathetic reader would care because the result redraws the line where AI systems stand relative to specialized human expertise in biology, with immediate consequences for biosecurity risk assessment and for how future biology benchmarks should be built.","feed_headline":"o3 doubles expert virologists on a biology test","feed_subtitle":"A 27-model, eight-benchmark sweep finds several models at or above expert baselines, while older tests saturate.","key_machinery":"The machinery is a standardized evaluation protocol: 27 models, each run ten times zero-shot on eight benchmarks, scored by exact answer-format matching, with performance summarized as mean accuracy and compared to random-guess and human-expert baselines. The central comparison device is a normalized score, $\\frac{\\text{model accuracy}-\\text{expert accuracy}}{1-\\text{expert accuracy}}$, which turns 'above expert' into a positive fraction of the expert-to-perfect gap. The two hardest benchmarks use formats that suppress guessing: VCT-Text is multiple-response, requiring selection of all true statements from 4-10 options, giving a random baseline near 2.5%, while CloningScenarios embeds multi-plasmid DNA sequences that the model must reason over without tools. This design lets the paper attribute gains to model capability rather than answer memorization, provided the external human baselines are comparable.","core_discovery":"The paper's central claim is that, as of early 2025, several frontier LLMs perform at or above the level of human experts on challenging biology question sets, not just on easier knowledge-recall tests. The strongest evidence is VCT-Text, where o3 scores 46.1% versus the 22.6% mean accuracy of PhD-level virologists on text-only questions, a more than doubling; on GPQA-Bio o3 reaches 81.2% against the 66.7% expert baseline; on WMDP-Bio Claude 3.7 Sonnet reaches 86.1% against the 60.5% expert baseline; and on CloningScenarios o3 and Claude 3.7 Sonnet score about 61% against the 60% expert baseline. The paper also claims that chain-of-thought prompting does not reliably improve zero-shot accuracy, while explicit reasoning-effort controls generally improve performance in line with inference scaling, with VCT-Text as the notable exception. These results are presented as evidence that the bottleneck in biology evaluation has shifted from model capability to benchmark design and human-baseline quality.","pith_inferences":["Editorial inference: if the expert baselines hold, expert-written multiple-choice questions no longer separate frontier models from top human specialists, so the measurable frontier moves to agentic tasks like literature retrieval, protocol execution, and experimental design.","Editorial inference: the paper's VCT reasoning anomaly suggests a testable fix: on exhaustive multiple-response formats, asking a model to first propose candidate answers and then prune them may recover accuracy lost when extra reasoning adds incorrect statements.","Editorial inference: because model scores on WMDP-Bio exceed the expert baseline, biosecurity evaluations that treat expert-level knowledge as the risk ceiling should be recalibrated upward.","Editorial inference: a held-out benchmark with time-stamped questions published after each model's training cutoff would distinguish genuine biological reasoning from memorization of benchmark content, which the present design cannot fully exclude."],"forward_implications":["If the central claim is correct, static multiple-choice biology benchmarks such as PubMedQA, MMLU-Bio, and WMDP-Bio are saturating below 100% because of benchmark limitations, so new evaluations with higher ceilings and better question validation are needed.","If correct, zero-shot evaluation is a reliable baseline for biology capability tracking, since five-shot and chain-of-thought prompting rarely improve accuracy and CoT costs 75x more output tokens on average.","If correct, frontier models at 61% on CloningScenarios can solve complex multi-fragment cloning design problems without the DNA manipulation tools that the expert baseliners were allowed to use.","If correct, o3's 46.1% on VCT-Text means an LLM can offer expert-level troubleshooting advice for wet-lab virology experiments, a capability with direct dual-use implications.","If correct, reasoning-effort controls improve accuracy on single-answer biology benchmarks, while on exhaustive multiple-response formats extra reasoning can reduce accuracy by inducing extra wrong selections."],"supporting_citations":[{"why":"Supplies the GPQA-Bio questions and the 66.7% PhD-expert and 43.2% non-expert baselines the paper compares against.","marker":"[22]"},{"why":"Supplies the VCT-Text questions and the 22.6% expert virologist baseline that anchors the paper's strongest claim.","marker":"[27]"},{"why":"Supplies the WMDP-Bio dataset, the biosecurity knowledge set whose expert comparison is central to the results.","marker":"[19]"},{"why":"Provides the 60.5% expert accuracy baseline for WMDP-Bio, measured on a sampled subset of questions.","marker":"[28]"},{"why":"Supplies the LitQA2, CloningScenarios, and ProtocolQA questions and their expert baselines of 70%, 60%, and 79%.","marker":"[18]"},{"why":"Supplies the PubMedQA test set and the single-annotator 78% human baseline used for that benchmark.","marker":"[25]"},{"why":"Supplies the MMLU-Bio test questions and the 89.8% estimated expert accuracy used as the expert line.","marker":"[26]"},{"why":"Documents high error rates in MMLU questions, supporting the paper's benchmark-saturation interpretation.","marker":"[37]"},{"why":"Provides the inference-scaling expectation that more reasoning tokens improve accuracy, against which the reasoning-effort results are read.","marker":"[31]"}],"fun_headline_variants":["o3 doubles expert virologists on hard text test","Frontier LLMs match or beat PhDs on biology benchmarks","o3 doubles expert virologists on VCT text questions","Biology benchmarks plateau as LLMs leap past experts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The human-expert baselines were measured on different, usually smaller, subsets of each benchmark than the full question sets the models answered, so the model-versus-expert comparisons may compare different tasks.","fun_headline_variants_meta":{"raw":{"variants":["o3 doubles expert virologists on hard text test","Frontier LLMs match or beat PhDs on biology benchmarks","o3 doubles expert virologists on VCT text questions","Biology benchmarks plateau as LLMs leap past experts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000527,"raw_usage":{"total_tokens":2564,"prompt_tokens":985,"completion_tokens":1579,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":1512}},"tokens_in":601,"tokens_out":1579,"duration_ms":11489,"temperature":1.0,"reasoning_tokens":1512,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:47:12.164927+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retest the 101 VCT-Text questions with a panel of PhD virologists under the same no-LLM conditions used in the original baselining, scoring exact multiple-response matches; if their mean accuracy reaches 46.1% or higher, the paper's headline claim that o3 performs twice as well as expert virologists fails on that benchmark.","supporting_citations":[],"review_version":1}