Pith. sign in

REVIEW 12 cited by

SciCode: A Research Coding Benchmark Curated by Scientists

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13168 v1 pith:YLIE675S submitted 2024-07-18 cs.AI cs.CL

classification cs.AIcs.CL
keywords scicodeproblemsscientificchallengingbenchmarkcodecodingevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we created a scientist-curated coding benchmark, SciCode. The problems in SciCode naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems. It offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. Claude3.5-Sonnet, the best-performing model among those tested, can solve only 4.6% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards becoming helpful scientific assistants and sheds light on the development and evaluation of scientific AI in the future.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agentic Exploration of Physics Models

    cs.AI 2025-09 conditional novelty 7.0 of 10

    A general-purpose LLM agent can discover physics models, including ODEs and spin Hamiltonians, by autonomously choosing experiments and fitting hypotheses to numeric data.

  2. CLVisc Agent for autonomous relativistic hydrodynamics studies

    nucl-th 2026-07 conditional novelty 6.0 of 10

    An LLM agent autonomously created a CLVisc skill and ran two hydrodynamic studies, finding that the high-temperature branch of η/s dominates flow suppression and that PGCM-uniform 16O decouples ellipticity from size.

  3. From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Procedure-guided teacher solutions let a 9B model gain durable SciCode skill that raw runtime procedures and matched no-procedure SFT do not provide.

  4. LLMoxie: Exploring Agentic AI for Scientific Software Development

    cs.SE 2026-07 conditional novelty 6.0 of 10

    A governed multi-cloud AI platform plus hierarchical RSE plugins turns generic coding agents into domain-aware collaborators for scientific software with provenance.

  5. ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new open-source benchmark evaluates LLM-generated end-to-end ML pipelines from Kaggle competition descriptions translated into 13 languages, with 6 private tasks to limit data leakage.

  6. Symmetry-induced magnetism in fullerene monolayers

    cond-mat.mtrl-sci 2025-08 unverdicted novelty 6.0 of 10

    The submitted paper could not be evaluated: the abstract describes symmetry-induced magnetism in fullerene monolayers, but the body text belongs to an unrelated LLM benchmark paper.

  7. Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LIMITGEN evaluates LLMs on identifying paper limitations and shows strong models still miss roughly half of obvious flaws, with retrieval augmentation giving modest but consistent gains.

  8. VASP Agent: An Agentic Framework for Autonomous First-principles Calculations

    cs.AI 2025-12 conditional novelty 5.0 of 10

    An LLM-driven agent with predefined VASP workflows and parameter-checking tools completes DFT simulation tasks more reliably and accurately than standalone LLMs, with a new 80-task benchmark.

  9. ChatVis: Large Language Model Agent for Generating Scientific Visualizations

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A retrieval-augmented LLM assistant with iterative error correction nearly doubles the rate of generating executable ParaView visualization scripts compared with unassisted models.

  10. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  11. How Far Are AI Scientists from Changing the World?

    cs.AI 2025-07 conditional novelty 4.0 of 10

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

  12. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.

Pith tools