Pith. sign in

REVIEW 3 major objections 5 minor 5 cited by

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read LLM-generated research ideas suffer far larger review-score drops than human expert ideas once they are actually executed, erasing the ideation-stage advantage AI ideas had shown.

desk verdict A first-of-its-kind execution study whose central gap result may be a regression-to-the-mean artifact rather than a real ideation-execution gap. read the letter →

arxiv 2506.20803 v1 pith:B4R6LZWI submitted 2025-06-25 cs.CL cs.AIcs.CYcs.HCcs.LG

classification cs.CLcs.AIcs.CYcs.HCcs.LG
keywords LLM-generatedresearchideasideation-executiongapideaevaluationexecutionstudyhumanrandomizedcontrolledtrialAIscientistsquality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether AI-generated research ideas that look good on paper still look good after real researchers spend months executing them. The authors took research ideas written by expert NLP researchers and by an LLM, randomly assigned 43 qualified researchers to execute them blind, and had 58 expert reviewers score the resulting projects. The central finding is an ideation–execution gap: on a 10-point scale, the LLM ideas' scores fell by roughly 1.0 to 2.0 points after execution on novelty, excitement, effectiveness, and overall quality, while human ideas barely moved, shrinking and on some metrics reversing AI's ideation-stage advantage. The paper argues that judging ideas without executing them systematically overstates the value of LLM-generated ideas, because only implementation exposes weaknesses such as infeasible plans, optimistic human-evaluation proposals, and unverified assumptions.

What carries the argument

The load-bearing object is the ideation–execution gap: for each idea, the difference between its post-execution review score and its pre-execution ideation score, computed on the four metrics shared by both evaluations (novelty, excitement, effectiveness, overall). Measuring the within-idea change from before to after execution removes much of the heterogeneity in idea quality that makes direct human-versus-AI comparisons statistically noisy, and this paired difference is what yields significant results at a sample size of 43. The comparison rests on a blinded, pre-registered randomized design: 43 expert executors each spent roughly 100 hours implementing a randomly assigned idea and wrote a 4-page paper, and 58 blinded expert reviewers produced 181 reviews, with the human and AI conditions indistinguishable on the control metrics of faithfulness and codebase quality.

What would settle it

Score a subsample of the same executed projects under the ideation protocol — reviewers see only the original idea text, without the paper or codebase — and check whether AI ideas' drop relative to human ideas persists when the evaluation protocol is held fixed; if the differential drop disappears, the gap is an artifact of the two evaluation regimes rather than of idea quality. The converse check is to have a second team of executors independently re-execute the same ideas: if the larger AI drop reproduces, the gap is robust.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLM-generated research ideas incur a significantly larger drop in expert review scores once they are executed than human expert ideas do. Using 43 executed projects (19 human, 24 AI) drawn from a prior blinded ideation study, the authors compute, for each idea, the difference between its execution review score and its ideation score on the four metrics both evaluations share. AI ideas drop by 1.049 to 1.976 points across novelty, excitement, effectiveness, and overall score; human ideas change by only −0.010 to −0.628, and the difference between the two conditions is significant on every metric (FDR-corrected p < 0.05). Because both conditions were executed under identical instructions, reviewed blind, and matched on faithfulness to the original idea and codebase quality, the authors attribute the differential drop to the origin of the idea. They further show that the drop is not explained by the executors' minor experiment-detail changes, and that execution reviewers weigh empirical performance, experimental rigor, and feasibility, factors that are almost invisible at the ideation stage.

Load-bearing premise

The most fragile premise is that the ideation-stage scores (from the earlier study) and the execution-stage scores (from this study) rate the same quality on the same numeric scale, so that a larger pre-post drop for AI ideas reflects worse ideas rather than a shift in reviewer standards, rubric anchors, or evaluation context between the two studies.

Editorial extensions

If this is right

  • Idea-stage reviews, by themselves, overstate the value of LLM-generated research ideas relative to human expert ideas; execution outcomes are needed to see their true quality.
  • The gap closes, and on excitement and effectiveness even reverses, the ranking between AI and human ideas observed at ideation, although the reversed ranking is not statistically significant at this sample size.
  • Ideation scores are weak predictors of execution outcomes, with correlations weak in most cases and moderately negative for AI ideas on excitement, so ideation evaluation should not be used as a proxy for research impact.
  • Evaluations of AI-scientist systems that stop at proposal or paper review without independent execution will likely overestimate these systems' research capability.
  • Reviewer rationales show that execution evaluation surfaces empirical performance, experiment rigor, and resource feasibility, factors that ideation evaluation cannot see because it implicitly assumes the proposed method will work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The result suggests a proxy-reward failure that likely generalizes beyond this study: LLMs may be optimizing ideas for surface signals that impress reviewers, such as novelty and excitement, rather than for executability, so any automated pipeline that scores ideas without executing them will drift toward ideas that look good and fail under implementation.
  • A testable extension: if the gap is driven by feasibility mismatches, stratifying AI ideas by the resources they propose (for example, planned human evaluations or heavy computation) should predict gap size; AI ideas that proposed human studies were indeed more likely to be altered by executors, and removing those six ideas did not change the result.
  • A crossover replication, in which the same ideas are executed by a second independent set of researchers, would separate idea quality from executor variance and reveal how much of the gap is attributable to the idea itself rather than its implementer.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a pre-registered execution study in which 43 expert researchers were randomly assigned to execute 19 human-written and 24 LLM-generated NLP research ideas taken from the authors' prior ideation study (Si et al., 2025). Each participant spent roughly 100 hours implementing the assigned idea and writing a 4-page paper, and 58 blind reviewers scored the resulting papers and codebases (181 reviews). The authors compare execution scores with the ideation scores from the prior study and find that AI ideas drop more between ideation and execution than human ideas on novelty, excitement, effectiveness, and overall (e.g., -1.049 vs -0.010 for novelty, Table 5). They also analyze executor changes, reviewer rationales, and robustness to removing six AI ideas whose human-evaluation components were dropped.

Significance. If the result holds, it is an important and timely caveat to claims that LLM-generated research ideas are superior to human ideas: idea quality judged without execution would overstate LLM idea value. The study's strengths include pre-registration, randomization and blinding, release of data, substantial per-project execution effort, and a useful robustness check. Its main quantitative claim, however, rests on a change-score comparison that is vulnerable to regression to the mean and to uncalibrated evaluation scales; the current analyses do not rule out these artifacts. The study is nonetheless a valuable and original contribution to the methodology of evaluating AI-generated research ideas, and the authors' qualitative analyses (Section 5) are informative.

major comments (3)
  1. [Section 4.2, Tables 4 and 5, Appendix D] The change-score analysis is confounded by regression to the mean. Because the outcome D = E - I is mechanically negatively correlated with I whenever I contains measurement error, and Table 9 reports near-zero ideation-execution correlations (AI novelty r = -0.019, excitement r = -0.321; human |r| <= 0.21), a group with higher mean ideation scores will tend to show larger negative mean drops even if its true execution outcomes are identical. Table 4 shows AI ideas start significantly higher at ideation (e.g., effectiveness 6.003 vs 4.833); a null model with execution scores independent of ideation would already predict gap differences of roughly 0.79-1.25 points, which is a large fraction of the observed Delta = 1.35-1.84 in Table 5. The statement in Section 4.2 that the gap 'controls for heterogeneity in idea quality' is therefore inaccurate for this design: difference scores remove stable idea-level effects only under strong assumptions that are not met here. Please re-analyze the data with ANCOVA on execution score adjusting for ideation score, or with within-ideation-strata matching, and report what the null model predicts for Table 5 before attributing the gap to condition.
  2. [Section 4.2 and Appendix A] The ideation and execution evaluations are not on a common, calibrated scale. Ideation reviewers scored short idea proposals, whereas execution reviewers scored 4-page papers and codebases using a different review form that adds soundness, codebase quality, and faithfulness, and the overall-score rubric is anchored to a 'short paper track at *ACL' (Appendix A). The gap D = execution score minus ideation score therefore conflates genuine idea-quality changes with any shift in reviewer standards, rubric anchors, or task demands between the two studies. Because the AI and human conditions have different ideation score distributions, even a common additive or multiplicative shift in scoring standards would produce differential observed gaps. The authors state that the review guidelines 'closely match' the ideation study, but this is not a calibration. Please provide evidence of scale comparability, for example through a crossover review of a common set of ideas or a sensitivity analysis allowing for study-specific shifts in score distributions, or temper the causal interpretation of the gap accordingly.
  3. [Section 2.1 and Table 2] The randomization of ideas to executors was within preferred topics, but the realized samples show large imbalances in potentially important covariates: execution participants in the Human condition spent more hours on average (112.6 vs 93.7) and reported lower topic familiarity (2.9 vs 3.4) than those in the AI condition. Since execution quality is partly a function of executor effort and familiarity, these imbalances could contribute to the observed execution-score differences and hence to the gap. Please report balance tests for these covariates and show that the Table 5 conclusions survive adjustment for time spent and familiarity, or at least discuss the sensitivity of the results to these imbalances.
minor comments (5)
  1. [Throughout] There are several typos and formatting errors: 'heterogenity' in Section 4.2, 'reponsibility' in Section 7.1, 'Learning positive' in the Appendix A excitement scale, and spacing issues in Table 4 (e.g., '4 .404').
  2. [Figure 2] The axis label '(Study2 Study1)' in Figure 2 is malformed; it should read something like '(Study 2 - Study 1)'.
  3. [Section 5.2, Figure 3] The manual categorization of reviewer rationales into the ten categories in Figure 3 would benefit from reporting inter-annotator agreement, since the categories are subjective and the figure underlies the qualitative claim that execution reviews consider more factors.
  4. [Table 3] The review-level t-tests in Table 3 treat each review as independent, which ignores clustering by idea and reviewer; the authors do report idea-level results, but the text should be explicit that the review-level p-values are exploratory.
  5. [Related Work] Wen et al. (2025) is cited in the Future Work section on proxy reward models but is not discussed in Related Work; a brief mention there would help contextualize the execution-outcome-prediction thread.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the execution scores are new measurements, and the ideation-execution gap is a comparison of two independent evaluations rather than a quantity derived from the paper's own assumptions.

full rationale

The paper's central claim is that LLM-generated ideas show a larger score drop than human ideas when comparing ideation-stage ratings (from Si et al., 2025) with execution-stage ratings collected in this study. The execution scores are newly collected blind reviews of executed projects; they are not computed from the ideation scores, and no parameter is fitted to the target gap. The reuse of the authors' own prior ideation study is a data dependency, not a circular derivation: the prior study is a published, externally available large-scale human evaluation, and the current paper's contribution is the new execution study. The difference-score analysis defines the gap as execution score minus ideation score, but this is a straightforward operationalization of the research question, not a self-referential construction. Statistical concerns about the comparability of the two rating scales or regression to the mean under baseline imbalance are legitimate methodological risks, but they concern the validity of causal inference, not circularity of the derivation. The paper also makes no appeal to a uniqueness theorem from the authors' prior work, and the main result does not reduce to any fitted parameter or to a self-citation chain.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters or invented entities are used in the central analysis. The paper's load-bearing assumptions are about the comparability and validity of the two evaluation studies, the balancing properties of the randomization, and the faithfulness of execution.

assumptions (5)
  • domain assumption Expert review scores on the provided rubric are valid and comparable measures of research idea and execution quality.
    Section 2.1 and Appendix A: the study treats reviewer scores on novelty, excitement, effectiveness, and overall as the outcome measure; no independent validation of the scale is given.
  • domain assumption Ideation-stage scores from the authors' prior study (Si et al., 2025) are on the same scale as execution-stage scores in this study.
    Section 4.2 computes gaps as execution score minus ideation score; if the two evaluations have different implicit standards, the gap is partly an artifact of the measurement mode.
  • domain assumption Random assignment of ideas to executors, conditional on topic preference, balances executor skill and effort across the two conditions.
    Section 2.2 and Table 2: AI-condition executors report higher topic familiarity but spend about 19 fewer hours on average; the paper does not adjust for these differences.
  • domain assumption The 43 completed projects are representative of the 66 onboarded participants, with no condition-specific attrition that biases the gap.
    Section 3.1 reports 23 non-completions, mostly personal reasons, but does not break down attrition by condition or identify the condition of the one idea excluded as too vague.
  • domain assumption Executors mostly preserved the original ideas, with changes limited to experimental details, so execution scores reflect the original ideas rather than the executors' modifications.
    Section 5.1 manually verifies change logs; this verification depends on the authors' judgment and cannot be independently checked from the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas." pith.science (2026). https://pith.science/paper/B4R6LZWI

@misc{pith2026250620803,
  author       = {Pith},
  title        = {Pith review of: The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B4R6LZWI}},
  note         = {Machine review of arXiv:2506.20803}
}
read the original abstract

Large Language Models (LLMs) have shown promise in accelerating the scientific research pipeline. A key capability for this process is the ability to generate novel research ideas, and prior studies have found settings in which LLM-generated research ideas were judged as more novel than human-expert ideas. However, a good idea should not simply appear to be novel, it should also result in better research after being executed. To test whether AI-generated ideas lead to better research outcomes, we conduct an execution study by recruiting 43 expert researchers to execute randomly-assigned ideas, either written by experts or generated by an LLM. Each expert spent over 100 hours implementing the idea and wrote a 4-page short paper to document the experiments. All the executed projects are then reviewed blindly by expert NLP researchers. Comparing the review scores of the same ideas before and after execution, the scores of the LLM-generated ideas decrease significantly more than expert-written ideas on all evaluation metrics (novelty, excitement, effectiveness, and overall; p < 0.05), closing the gap between LLM and human ideas observed at the ideation stage. When comparing the aggregated review scores from the execution study, we even observe that for many metrics there is a flip in rankings where human ideas score higher than LLM ideas. This ideation-execution gap highlights the limitations of current LLMs in generating truly effective research ideas and the challenge of evaluating research ideas in the absence of execution outcomes.

Figures

Figures reproduced from arXiv: 2506.20803 by the authors.

Figure 1
Figure 1. Study overview: we recruit N = 43 expert researchers to execute randomly assigned ideas from either the Human condition or the AI condition. Expert reviewers then blindly review all the executed projects. Despite the AI ideas being scored higher than human ideas before execution (e.g., their predicted effectiveness score of the ideas), their scores drop significantly more than human ideas after execution (e.g., thei… view at source ↗
Figure 2
Figure 2. Average scores of AI ideas drop significantly more than Human ideas in the execution study [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Comparison of the factors mentioned in the reviewer rationales in the ideation (yellow [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (25 more)
Figure 4
Figure 4. Figure 4: Boxplots of all evaluation metrics across Human and AI conditions in the execution evalua [PITH_FULL_IMAGE:figures/full_fig_p025_4.png]
Figure 1
Figure 1. Figure 1: Workflow of our proposed compound LLM system for mimicing knowledge unlearning. [PITH_FULL_IMAGE:figures/full_fig_p035_1.png]
Figure 1
Figure 1. Figure 1: Self-Improving Memory Framework. A math problem is decomposed into subtasks, solved inde￾pendently, and stored in a static memory pool with embeddings. At test time, new subtasks retrieve similar past examples to guide memory-augmented inference. 2.1 Subtask Decomposit…
Figure 1
Figure 1. Figure 1: Reliability diagram comparing calibration of [PITH_FULL_IMAGE:figures/full_fig_p069_1.png]
Figure 2
Figure 2. Figure 2: Reliability diagram comparing different cali [PITH_FULL_IMAGE:figures/full_fig_p070_2.png]
Figure 3
Figure 3. Figure 3: Reliability diagram comparing different cali [PITH_FULL_IMAGE:figures/full_fig_p071_3.png]
Figure 1
Figure 1. Figure 1: MKA Pipeline confusion matrix These met￾rics consider abstentions and responses indi￾vidually. These accuracies there￾fore do not allow meaningful com￾parison: the ratio of abstentions￾responses changes de￾pending on the confidence cutoff, which causes the denominators…
Figure 2
Figure 2. Figure 2: Accuracy analyses for the MKA Pipeline on Aya Expanse 8B, Gemma2 9B and Qwen2.5 7B across the target languages [PITH_FULL_IMAGE:figures/full_fig_p080_2.png]
Figure 1
Figure 1. Figure 1: Left and Center: Accuracy and demographic parity difference on BiosBias dataset with one pivot injected per prompt, broken down by prompt ID. Pivot ID 0 corresponds to the base prompt (no pivot). Right: Idealized CAT score on StereoSet dataset with one pivot injected p…
Figure 2
Figure 2. Figure 2: Left: Accuracy and demographic parity difference on BiosBias dataset as functions of the number of pivots injected per prompt. Zero pivots corresponds to the base prompt. Right: Language modeling score, stereotype score, and idealized CAT score on StereoSet dataset as …
Figure 1
Figure 1. Figure 1: ASR(%) on Vicuna-7B for harmful prompts. [PITH_FULL_IMAGE:figures/full_fig_p114_1.png]
Figure 2
Figure 2. Figure 2: Example where GPT-4o complies with masked prompt. Dataset ASR(%) in Target LLMs GPT-4o Llama3-8b Llama2-7b Vicuna-7b Adv-Bench Unmasked 0 0.38 0.38 33.08 NAM 0 0.58 0.96 34.80 ASM (ours) 10 2.69 2.12 13.08 JBB-harmful Unmasked 0 3 0 42 NAM 0 3 3 44 ASM (ours) 3 4 2 9 J…
Figure 3
Figure 3. Figure 3: ASR (%) for our optimized attack on GPT-4o. [PITH_FULL_IMAGE:figures/full_fig_p115_3.png]
Figure 4
Figure 4. Figure 4: Prompt for category generation by GPT-4o [PITH_FULL_IMAGE:figures/full_fig_p118_4.png]
Figure 5
Figure 5. Figure 5: Prompt for span highlighting by GPT-4o A.2 Choice of LLM for Masking During our dataset generation process, we experimented with different LLMs for category generation and span highlighting. With Vicuna, the response quality and formatting were comparatively inconsiste…
Figure 6
Figure 6. Figure 6: Prompt for evaluation by the judge model (GPT-4o) [PITH_FULL_IMAGE:figures/full_fig_p121_6.png]
Figure 7
Figure 7. Figure 7: ASR(%) on different LLMs for harmful prompts. [PITH_FULL_IMAGE:figures/full_fig_p122_7.png]
Figure 8
Figure 8. Figure 8: GPT-4o complying with harmful request from Adv-Bench. [PITH_FULL_IMAGE:figures/full_fig_p123_8.png]
Figure 9
Figure 9. Figure 9: Llama2-7B complying with harmful request from Adv-Bench. [PITH_FULL_IMAGE:figures/full_fig_p124_9.png]
Figure 10
Figure 10. Figure 10: Llama3-8B complying with harmful request from Adv-Bench. [PITH_FULL_IMAGE:figures/full_fig_p125_10.png]
Figure 11
Figure 11. Figure 11: Vicuna complying with harmful request from Adv-Bench. [PITH_FULL_IMAGE:figures/full_fig_p126_11.png]
Figure 12
Figure 12. Figure 12: GPT-4o complying with harmful request from JBB-harmful [PITH_FULL_IMAGE:figures/full_fig_p127_12.png]
Figure 13
Figure 13. Figure 13: GPT-4o complying with an optimized-harmful request from Adv-Bench. [PITH_FULL_IMAGE:figures/full_fig_p128_13.png]
Figure 14
Figure 14. Figure 14: Vicuna providing a harmless response to a harmful request from DAN-fq. [PITH_FULL_IMAGE:figures/full_fig_p129_14.png]
Figure 15
Figure 15. Figure 15: Vicuna providing a harmless response to a harmful request from JBB-harmful. [PITH_FULL_IMAGE:figures/full_fig_p130_15.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Conference accept/reject outcomes yield 15 operational ideation patterns that, as an LLM skill suite, improve automated-judged research-proposal quality over no-skill and generic-skill baselines.

  2. Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    LLM judges of scientific ideas are measurably swayed by writing style; a style-detecting module reduces but does not remove the bias.

  3. Budgeted Subset Refinement for Execution-Aware LLM Research Ideation

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Refining a diversity-selected subset (MMR-k) of LLM-generated research ideas yields the best proxy-rated portfolio quality per compute, while raw ideas and reranking alone fail to produce strong nonduplicate proposals.

  4. AI Can Learn Scientific Taste

    cs.CL 2026-03 conditional novelty 6.0 of 10

    Reinforcement learning on citation-preference pairs teaches a model to predict which papers will be cited more and to propose ideas that LLM judges rate as likely to be cited more—but "taste" here means citation impact.

  5. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

Reference graph

Works this paper leans on

31 extracted references · 11 canonical work pages · cited by 5 Pith papers

  1. [1]

    ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models

    Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models . In NAACL, 2025

  2. [2]

    Bartoldson, Bhavya Kailkhura, Tom Goldstein, and Furong Huang

    Zikui Cai, Shayan Shabihi, Bang An, Zora Che, Brian R. Bartoldson, Bhavya Kailkhura, Tom Goldstein, and Furong Huang. AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security . ICLR Workshop BuildingTrust, 2025. URL https://arxiv.org/abs/2504.20965

  3. [3]

    MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

    Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Lilian Weng, and Aleksander Mkadry. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering . In ICLR, 2025

  4. [4]

    MKA: Leveraging Cross-Lingual Consensus for Model Abstention

    Sharad Duwal. MKA: Leveraging Cross-Lingual Consensus for Model Abstention . ICLR Workshop BuildingTrust, 2025. URL https://arxiv.org/abs/2503.23687

  5. [5]

    GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation

    Tao Feng, Yihang Sun, and Jiaxuan You. GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation . In ICLR, 2025

  6. [6]

    Szostkiewicz, Jon M

    Ali Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J. Szostkiewicz, Jon M. Laurent, Muhammed T. Razzak, Andrew D. White, Michaela M. Hinks, and Samuel Rodriques. Robin: A multi-agent system for automating scientific discovery . ArXiv, 2025. URL https://arxiv.org/abs/2505.13400

  7. [7]

    Towards an AI co-scientist

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, Yossi Matias, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Yuan Guan, Vi...

  8. [8]

    Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas

    Xiang Hu, Hongyu Fu, Jinge Wang, Yifeng Wang, Zhikun Li, Renjun Xu, Yu Lu, Yaochu Jin, Lili Pan, and Zhenzhong Lan. Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas . ArXiv, 2024. URL https://arxiv.org/abs/2410.14255

Show all 31 references
  1. [9]

    Truong, Weixin Liang, Fan-Yun Sun, and Nick Haber

    Tianyu Hua, Harper Hua, Violet Xiang, Benjamin Klieger, Sang T. Truong, Weixin Liang, Fan-Yun Sun, and Nick Haber. ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code . arXiv, 2025. URL https://arxiv.org/abs/2506.02314

  2. [10]

    Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents

    Long Li, Weiwen Xu, Jiayan Guo, Ruochen Zhao, Xinxuan Li, Yuqian Yuan, Boqiang Zhang, Yuming Jiang, Yifei Xin, Ronghao Dang, Deli Zhao, Yu Rong, Tian Feng, and Li Bing. Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents . In ICLR, 2025

  3. [11]

    ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering

    Zexi Liu, Jingyi Chai, Xinyu Zhu, Shuo Tang, Rui Ye, Bo Zhang, Lei Bai, and Siheng Chen. ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering . arXiv, 2025. URL https://arxiv.org/abs/2505.23723

  4. [12]

    The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

    Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery . ArXiv, 2024. URL https://arxiv.org/abs/2408.06292

  5. [13]

    Kozlovskii, Francisco J

    Alexander Novikov, Ng \^a n V˜u, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav M. Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin...

  6. [14]

    Sparks of Science: Hypothesis Generation Using Structured Paper Data

    Charles O'Neill, Tirthankar Ghosal, Roberta Raileanu, Mike Walmsley, Thang Bui, Kevin Schawinski, and Ioana Ciuca. Sparks of Science: Hypothesis Generation Using Structured Paper Data . ArXiv, abs/2504.12976, 2025. URL https://arxiv.org/abs/2504.12976

  7. [15]

    PolyPrompt: Automating Knowledge Extraction from Multilingual Language Models with Dynamic Prompt Generation

    Nathan Roll. PolyPrompt: Automating Knowledge Extraction from Multilingual Language Models with Dynamic Prompt Generation . ArXiv, abs/2502.19756, 2025. URL https://arxiv.org/abs/2502.19756

  8. [16]

    Agent Laboratory: Using LLM Agents as Research Assistants

    Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum. Agent Laboratory: Using LLM Agents as Research Assistants . ArXiv, abs/2501.04227, 2025. URL https://arxiv.org/abs/2501.04227

  9. [17]

    Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

    Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers . In ICLR, 2025. URL https://arxiv.org/abs/2409.04109

  10. [18]

    Do grant proposal texts matter for funding decisions? a field experiment

    M \"u ge Simsek, Mathijs de Vaan, and Arnout van de Rijt. Do grant proposal texts matter for funding decisions? a field experiment. Scientometrics, 129: 0 2521--2532, 2024. URL https://link.springer.com/article/10.1007/s11192-024-04968-7

  11. [19]

    PaperBench: Evaluating AI's Ability to Replicate AI Research

    Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. PaperBench: Evaluating AI's Ability to Replicate AI Research . ArXiv, 2025. ...

  12. [20]

    Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System

    Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, and Nanqing Dong. Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System . In ACL, 2025

  13. [21]

    SciMON: Scientific Inspiration Machines Optimized for Novelty

    Qingyun Wang, Doug Downey, Heng Ji, and Tom Hope. SciMON: Scientific Inspiration Machines Optimized for Novelty . In ACL, 2024 a

  14. [22]

    SciPIP: An LLM-based Scientific Paper Idea Proposer

    Wenxiao Wang, Lihui Gu, Liye Zhang, Yunxiang Luo, Yi Dai, Chen Shen, Liang Xie, Binbin Lin, Xiaofei He, and Jieping Ye. SciPIP: An LLM-based Scientific Paper Idea Proposer . ArXiv, 2024 b . URL https://arxiv.org/abs/2410.23166

  15. [23]

    Predicting Empirical AI Research Outcomes with Language Models

    Jiaxin Wen, Chenglei Si, Yueh han Chen, He He, and Shi Feng. Predicting Empirical AI Research Outcomes with Language Models . ArXiv, 2025. URL https://arxiv.org/abs/2506.00794

  16. [24]

    CycleResearcher: Improving Automated Research via Automated Review

    Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. CycleResearcher: Improving Automated Research via Automated Review . In ICLR, 2025

  17. [25]

    The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

    Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search . ArXiv, 2025. URL https://arxiv.org/abs/2504.08066

  18. [26]

    Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and E. Cambria. Large Language Models for Automated Open-domain Scientific Hypotheses Discovery . ACL Findings, 2024

  19. [27]

    Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

    Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, and Jeff Clune. Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents . ArXiv, 2025. URL https://arxiv.org/abs/2505.22954

  20. [28]

    DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process

    Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang. DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process . ArXiv, abs/2503.08569, 2025. URL https://arxiv.org/abs/2503.08569

  21. [29]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  22. [30]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  23. [31]

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.