Pith. sign in

REVIEW 6 cited by

Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.16069 v2 pith:MJF7K2VG submitted 2025-02-22 cs.AI cs.LG

classification cs.AIcs.LG
keywords curieexperimentationrigormodulescientificautomatingcontrolenhance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Scientific experimentation, a cornerstone of human progress, demands rigor in reliability, methodical control, and interpretability to yield meaningful results. Despite the growing capabilities of large language models (LLMs) in automating different aspects of the scientific process, automating rigorous experimentation remains a significant challenge. To address this gap, we propose Curie, an AI agent framework designed to embed rigor into the experimentation process through three key components: an intra-agent rigor module to enhance reliability, an inter-agent rigor module to maintain methodical control, and an experiment knowledge module to enhance interpretability. To evaluate Curie, we design a novel experimental benchmark composed of 46 questions across four computer science domains, derived from influential research papers, and widely adopted open-source projects. Compared to the strongest baseline tested, we achieve a 3.4$\times$ improvement in correctly answering experimental questions. Curie is open-sourced at https://github.com/Just-Curieous/Curie.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

    cs.AI 2026-04 reject novelty 6.0 of 10

    In a benchmark of four AI Scientist systems on 15 FARS proposals, LLM reviewers rated FARS's own papers about twice as high as Sakana v1/v2, CycleResearcher, and Data-to-Paper outputs, but the evaluation lacks human v...

  2. Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance

    cs.AI 2026-01 conditional novelty 6.0 of 10

    A multi-agent 'verify-then-write' system for writing peer-review rebuttals beats direct LLM prompting on a new benchmark, but the gains are measured by an LLM judge, not by the original reviewers.

  3. Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture

    cs.SE 2025-11 conditional novelty 6.0 of 10

    An AI-Scientist guard architecture combining a Haskell monad for online FDR accounting with declarative scaffolding against data leakage; simulation supports it, but the advertised Lean/SPARK verification is absent fr...

  4. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  5. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

  6. AI Scientists Fail Without Strong Implementation Capability

    cs.AI 2025-06 conditional novelty 4.0 of 10

    AI scientist systems can propose ideas but cannot reliably implement and verify experiments, making the implementation gap, not idea generation, the current bottleneck.

Pith tools