Pith. sign in

REVIEW 6 cited by

Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.16137 v2 pith:O6TY3QG3 submitted 2025-04-21 cs.CY cs.LG

classification cs.CYcs.LG
keywords virologyvirologistsdual-useexpertbenchmarkcapabilitiescapabilityexpert-level
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Constructed from the inputs of dozens of PhD-level expert virologists, VCT consists of $322$ multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories. VCT is difficult: expert virologists with access to the internet score an average of $22.1\%$ on questions specifically in their sub-areas of expertise. However, the most performant LLM, OpenAI's o3, reaches $43.8\%$ accuracy, outperforming $94\%$ of expert virologists even within their sub-areas of specialization. The ability to provide expert-level virology troubleshooting is inherently dual-use: it is useful for beneficial research, but it can also be misused. Therefore, the fact that publicly available models outperform virologists on VCT raises pressing governance considerations. We propose that the capability of LLMs to provide expert-level troubleshooting of dual-use virology work should be integrated into existing frameworks for handling dual-use technologies in the life sciences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI Security Leaderboard: Methodology, Results and Minimal Standard

    cs.CR 2026-08 conditional novelty 7.0 of 10

    A new framework, the FAR.AI Minimal Standard, measures frontier safeguards and finds Grok 4.5 and Gemini 3.1 Pro are cheaply jailbroken while Claude Fable 5 and GPT-5.6 Sol showed no universal jailbreaks under the sam...

  2. BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment

    cs.CR 2026-07 conditional novelty 6.5 of 10

    Across 16 model-harness setups, AI agents refuse legitimate literature-derived biology tasks at rates comparable to or higher than concealed biosecurity hazards, with most refusals coming from pre-reasoning API filters.

  3. BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

    cs.CY 2026-07 conditional novelty 6.0 of 10

    BioTIER, a 542-prompt benchmark with three risk tiers, shows frontier AI models differ by 90 percentage points in refusing dangerous biological queries, with top refusers over-refusing benign topics at the boundary.

  4. FORTRESS: Frontier Risk Evaluation for National Security and Public Safety

    cs.CY 2025-06 conditional novelty 6.0 of 10

    A new benchmark with instance-specific rubrics measures frontier LLMs' willingness to assist with national security and public safety threats, alongside a paired over-refusal test.

  5. An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

    cs.CL 2026-07 reject novelty 5.0 of 10

    A bio-red-teaming model is reported to jailbreak 14 frontier LLMs into producing dangerous biosecurity outputs, but the claimed wet-lab physical verification was not actually carried out.

  6. Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

    cs.AI 2025-07 conditional novelty 5.0 of 10

    An evaluation of 18 frontier AI models across seven catastrophic-risk categories finds all models in green or yellow zones, with none crossing the report's proposed red lines.

Pith tools