REVIEW 3 major objections 6 minor 63 references
Public ML conference outcomes contain reusable ideation patterns that can be turned into skills for writing stronger, evidence-grounded research proposals.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 19:10 UTC pith:6MIIPCJT
load-bearing objection Solid systems paper: contrastive Oral/HC/Reject pattern cards plus a real first-mile skill suite; the quality win is real inside its automated protocol but still partly style-aligned with the judge. the 3 major comments →
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The authors claim that Oral, high-citation, and rejected ML conference papers share a compact strategy space of about 15 reusable ideation patterns, that rejected work is best read as contrastive evidence about weak instantiations rather than a separate strategy class, and that packaging those patterns as structured cards inside a multi-phase, retrieval-and-audit skill produces research proposals that automated judges rate higher in quality than no-skill or generic-skill baselines while remaining competitively novel.
What carries the argument
The ideation-pattern card: a structured, corpus-derived object that pairs operational signature, when-to-apply conditions, success conditions from accepted work, failure modes from rejected work, and reviewer expectations—used both to select a research move for a diagnosed bottleneck and to audit the generated candidate against the same corpus-derived failure modes.
Load-bearing premise
That two automated LLM judging skills for idea quality and prior-art collision are close enough to human expert judgment to support the endpoint quality claim.
What would settle it
A blind human-expert ranking study on the same 100 method-agnostic ICLR 2026 seeds in which IdeaSpark no longer wins on quality relative to the same baselines, or in which human novelty judgments reverse the competitive-novelty result.
If this is right
- Early-stage research ideation can be treated as an evidence-grounded skill layer rather than open-ended brainstorming.
- Accepted and rejected ML papers often use the same high-level moves; the useful signal is execution quality and failure modes, not strategy choice alone.
- Pattern composition of one to three moves, with two as the modal default, better matches real papers than single-move recipes.
- Standalone literature-search and prior-art-collision skills can be reused outside the full ideation loop.
- Automated quality–novelty planes can expose “novel-but-empty” proposals that score high on novelty only because they are too vague to collide with prior art.
Where Pith is reading between the lines
- The same accept/reject contrastive-card idea could transfer to other fields that publish decisions and reviews, not only mainstream ML conferences.
- If human experts confirm the automated quality ranking, labs could use pattern cards as pre-experiment checklists rather than only as generation prompts.
- The finding that rejected constructive moves often drift toward audit/unify surface forms suggests a practical diagnostic for early drafts: check whether a construction has collapsed into a weaker audit claim.
- Separating generation freedom from audit anchoring may be a general design rule for other LLM research tools that currently over-constrain generation with corpus frequency priors.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ResearchStudio-Idea, a three-skill suite (Paper-Search, Scoop-Check, IdeaSpark) for the first mile of ML research ideation. From 1,947 ICLR/ICML/NeurIPS papers (2021–2025) labeled Oral, high-citation, and Reject, the authors extract domain-agnostic innovation signatures, cluster them into 31 sub-patterns, and induce 15 operational ideation-pattern cards that encode success conditions, failure modes, and differentiation strategies. IdeaSpark composes literature grounding, bottleneck diagnosis, pattern-guided candidate construction, prior-art collision retrieval, and outcome-informed audit into a single idea-card workflow with deterministic validators. A blind automated-judge study on 100 method-agnostic ICLR 2026 Oral seeds reports that IdeaSpark attains mean quality 3.87/4 (first on 88/100 seeds) versus same-backbone bare and generic-skill baselines, while remaining competitively novel on a scoop-check scale.
Significance. If the results hold, the work supplies a concrete missing layer between conference outcome data and inference-time ideation: reusable, failure-aware pattern cards induced from Oral/HC/Reject contrast rather than accepted-only taxonomies or parametric policies. Strengths include the three-way outcome contrast, reject-only remapping showing shared strategy space (§10), multi-label composition analysis with k=2 as mode (§6.5), embedding/abstraction ablations (§11), same-backbone evaluation ladder with length/format normalization (§14), and explicit faithfulness machinery (kill-switch fields, citation gates, honest abandon paths). The quality–novelty plane analysis that exposes a “novel-but-empty” failure mode is a useful methodological contribution beyond the system itself. The main open question is whether automated-judge gains transfer to human expert preference and downstream research value.
major comments (3)
- [§14 Evaluation] §14.4–§14.7 and Tables 11–12: The endpoint quality claim rests on an idea-quality judge that rewards problem position, method depth (dominating), soundness/feasibility, and problem-fit, with formal equations required by the normalization contract (§14.3). IdeaSpark is explicitly engineered to emit that surface (bottleneck diagnosis, sub-pattern instantiation, equations with interpretation, falsification/kill-switch, implementability audit; §13.3–13.4, App. A). The same-backbone ladder isolates skill design for generation, but both judges are LLM skills and no control rewrites IdeaSpark cards into bare-style prose (or expands bare outputs into IdeaSpark structure) before re-ranking. Without that control, cross-family judges, or a human pilot, the 3.87/4 and 88/100 wins may partly measure preference for the skill’s own output contract rather than independent research-move quality—the feasi
- [§6 Ideation-Pattern Induction] §1.4 and §6: The 15-pattern taxonomy is produced by a single Opus 4.7 induction call with a 6–18 size band; the paper notes that a comparable run would likely land at 12–18 with substantial overlap, but reports no inter-prompt, inter-seed, or inter-model stability study. Because the cards are the shared object for both generation and audit, instability in pattern boundaries would affect both Phase 2 selection and Phase 3 failure-mode checks. A modest multi-seed or multi-model remapping (e.g., Jaccard of cluster→pattern assignments, or card-level content agreement on a held-out paper sample) is needed to support the claim that the induced map is a reusable operational substrate rather than a single-run artifact.
- [§14.2 Problem seeds / §15 Limitations] §14.2 and §15: Evaluation seeds are method-agnostic rewrites of ICLR 2026 Oral titles. While held-out relative to the 2021–2025 induction corpus, they still sample the same venue’s Oral distribution and primary-area mix (Fig. 1 right). Combined with corpus restriction to three ML conferences, this leaves open whether quality gains hold for non-Oral problem framings, systems/benchmark-style directions that the quality rubric systematically down-weights (§15), or domains outside the induced 28-area map. The paper should either (i) add a non-Oral or cross-venue seed set, or (ii) more tightly scope the endpoint claim to “Oral-like ML conference directions under automated judges.”
minor comments (6)
- [Abstract / §1.4] Abstract and §1.4: Soften absolute phrasing (“consistently produces stronger research proposals”) to match the paper’s own scoping as an automated-judge, idea-stage study until human or style-control evidence is added.
- [§5.2 Clustering] §5.2 Table 2: 47.7% unclustered at the selected min_cluster_size is high; the multi-label pass (§6.5) addresses coverage, but a short quantitative comparison of multi-label pattern distributions for clustered vs unclustered papers would reassure readers that unclustered points are not systematically different strategies.
- [§7 Acceptance and Impact Analysis] §7 Table 7: ∆_OR uses cluster-level primaries only (N_O=482, N_R=367). State more prominently in the table caption that these are not full-class denominators, to avoid over-reading the tight ±2.9 pp spread.
- [Figure 1] Figure 1 left: Axis labels and baseline names are dense; a legend mapping “Opus-4.8 (self-gen)” to the generic auto-authored skill would help readers who land on the figure first.
- [§3.1 Scope and labeling] §3.1: Clarify how the 49 Oral∩HC overlaps are handled in every analysis that reports mutually exclusive classes versus inclusive HC_all, ideally with a single convention table.
- [Appendix] App. A–D are valuable; ensure the main text points to which appendix artifact corresponds to which evaluation seed ID so readers can audit a full run trail.
Circularity Check
No load-bearing circular derivation; mild residual that automated quality axes align with IdeaSpark’s engineered card contract, which the paper itself scopes as an automated-judge study.
full rationale
This paper does not claim a first-principles mathematical prediction that reduces to its inputs by construction. The empirical chain is corpus construction → strategy-signature extraction → clustering → 15/31 pattern cards → inference-time skill, then a separate held-out endpoint evaluation on method-agnostic ICLR 2026 Oral seeds. Evaluation seeds post-date the 2021–2025 induction corpus and strip each paper’s own method, so the quality/novelty results are not rediscovery of the training set. Pattern cards dual-used for generation and audit are an intentional design choice, not a self-definitional proof. Scoop-Check appears both as a suite skill and as the novelty judge, but novelty scores are not fitted targets that force IdeaSpark’s generation, and IdeaSpark is not the novelty leader (GPT-5.5 bare is). Format normalization (matched Title/Motivation/Method lengths; required formal equations for all sources) and the same-backbone ladder (Opus bare ≈ Opus-self-gen ≪ IdeaSpark) further block a pure “format equals quality” reduction. The residual concern—that idea-quality’s depth/problem-fit axes philosophically match what IdeaSpark is built to emit—is a construct-validity / human-alignment limitation the paper already states in §14–§15, not a circular step of the kinds this pass flags. No self-citation uniqueness theorem, fitted-as-prediction, or renamed known law anchors the central claim.
Axiom & Free-Parameter Ledger
free parameters (5)
- HDBSCAN min_cluster_size
- UMAP n_neighbors / n_components / min_dist
- Target taxonomy size band (6–18 patterns)
- Oral-safe / Reject-warn thresholds (p_O ≥65% / ≤35%)
- Idea-quality and scoop-check scoring rubrics
axioms (5)
- domain assumption Public ML conference outcomes (Oral, high-citation, Reject) contain reusable traces of how impactful research directions are formulated, differentiated, and evaluated.
- domain assumption Domain-agnostic rewriting of innovation fields yields strategy clusters rather than topic clusters.
- domain assumption LLM extraction, multi-label tagging, and pattern-card induction preserve operationally useful strategy structure.
- domain assumption Automated idea-quality and scoop-check judges provide a useful first-stage ranking of proposal quality and novelty.
- standard math Standard embedding/clustering mathematics (cosine similarity, UMAP, HDBSCAN, silhouette) is a valid substrate for grouping research strategies.
invented entities (3)
-
15 ideation-pattern cards / 31 sub-pattern cards
no independent evidence
-
IdeaSpark multi-phase skill (Phases 0–4 with kill-switch validators)
no independent evidence
-
Scoop-Check four-axis novelty collision level
no independent evidence
read the original abstract
Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating candidate directions. Researchers must ground a problem in current literature, identify meaningful bottlenecks, differentiate from existing solutions, and evaluate risks before committing to implementation. We present ResearchStudio-Idea as a reusable skill suite for this first mile of research ideation. The suite includes Paper-Search, a standalone multi-source literature search skill; Scoop-Check, a standalone prior-art collision checker for novelty claims; and IdeaSpark, the end-to-end skill that composes evidence grounding, pattern-guided generation, collision retrieval, audit, and idea-card rendering into one workflow. IdeaSpark is constructed from a corpus of 1,947 machine learning conference papers collected from ICLR, ICML, and NeurIPS between 2021 and 2025, including Oral papers, a separately tracked high-citation subset, and rejected submissions. Analysis of these outcomes reveals 31 recurring ideation sub-patterns, consolidated into 15 reusable ideation patterns. Each pattern is operationalized as a structured card containing research contexts, bottleneck types, differentiation strategies, supporting precedents, and common failure modes. Given a research problem and an evidence bundle, IdeaSpark evaluates evidence readiness, reconstructs the surrounding research context, identifies unresolved bottlenecks, selects relevant patterns, instantiates one candidate direction, retrieves potentially conflicting prior work, and performs outcome-informed auditing. This workflow transforms reusable ideation patterns into traceable research proposals. Blind automated-judge evaluations show that IdeaSpark consistently produces stronger research proposals than no-skill and generic-skill baselines while maintaining competitive novelty.
Figures
Reference graph
Works this paper leans on
-
[1]
SPECTER2 model card, 2026
Allen Institute for AI. SPECTER2 model card, 2026. Official SPECTER2 Hugging Face model card; checked May 10, 2026
2026
-
[2]
Claude models overview, 2026
Anthropic. Claude models overview, 2026. Official Anthropic model documentation; checked May 10, 2026
2026
-
[3]
Skill authoring best practices, 2026
Anthropic. Skill authoring best practices, 2026. Official Anthropic documentation; checked May 10, 2026
2026
-
[4]
Artiles, Martin Weiss, Levin Brinkmann, Anirudh Goyal, and Nasim Rahaman
Alejandro H. Artiles, Martin Weiss, Levin Brinkmann, Anirudh Goyal, and Nasim Rahaman. Alien science: Sampling coherent but cognitively unavailable research directions from idea atoms, 2026. arXiv:2603.01092; checked May 10, 2026
Pith/arXiv arXiv 2026
-
[5]
Agentic AI scientists are not built for autonomous scientific discovery, 2026
Harshit Bisht et al. Agentic AI scientists are not built for autonomous scientific discovery, 2026
2026
-
[6]
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. SPECTER: Document-level representation learning using citation-informed transformers. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2270–2282, 2020. ACL 2020; arXiv:2004.07180; checked May 10, 2026
Pith/arXiv arXiv 2020
-
[7]
Graphmind: Interactive novelty assessment system for accelerating scientific discovery, 2025
Italo Luis da Silva, Hanqi Yan, Lin Gui, and Yulan He. Graphmind: Interactive novelty assessment system for accelerating scientific discovery, 2025. arXiv:2510.15706; checked May 10, 2026
arXiv 2025
-
[8]
Graphs of research: Citation evolution graphs as supervision for research idea gener- ation, 2026
Songyang Gao et al. Graphs of research: Citation evolution graphs as supervision for research idea gener- ation, 2026
2026
-
[9]
Iris: Interactive research ideation system for accelerating scientific discovery, 2025
Aniketh Garikaparthi, Manasi Patwardhan, Lovekesh Vig, and Arman Cohan. Iris: Interactive research ideation system for accelerating scientific discovery, 2025. arXiv:2504.16728; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[10]
EigenGame: PCA as a nash equilibrium
Ian Gemp, Brian McWilliams, Claire Vernade, and Thore Graepel. EigenGame: PCA as a nash equilibrium. InInternational Conference on Learning Representations, 2021. ICLR 2021 Outstanding Paper; corpus paper id: ICLR 2021 0018; checked May 10, 2026
2021
-
[11]
Towards an AI co-scientist,
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, et al. Towards an AI co-scientist,
-
[12]
arXiv:2502.18864; checked May 10, 2026
Pith/arXiv arXiv 2026
-
[13]
BERTopic: Neural topic modeling with a class-based tf-idf procedure, 2022
Maarten Grootendorst. BERTopic: Neural topic modeling with a class-based tf-idf procedure, 2022. arXiv:2203.05794; checked May 10, 2026
Pith/arXiv arXiv 2022
-
[14]
Chenyang Gu, Jiahao Cheng, Meicong Zhang, Pujun Zheng, Jinquan Zheng, and Guoxiu He. Mori: Learn- ing motivation-grounded reasoning for scientific ideation in large language models, 2026. arXiv:2603.19044; checked May 10, 2026
Pith/arXiv arXiv 2026
-
[15]
Xuemei Gu and Mario Krenn. Interesting scientific idea generation using knowledge graphs and LLMs: Evaluations with 100 research group leaders, 2024. arXiv:2405.17044; checked May 10, 2026
Pith/arXiv arXiv 2024
-
[16]
Ideabench: Benchmarking large language models for research idea generation, 2024
Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, Eric Xie, Stefan Bekiranov, and Aidong Zhang. Ideabench: Benchmarking large language models for research idea generation, 2024. arXiv:2411.02429; checked May 10, 2026
Pith/arXiv arXiv 2024
-
[17]
The price of differential privacy under continual observation
Monika Henzinger, Satish Sricharan, and Teresa Steiger. The price of differential privacy under continual observation. InInternational Conference on Machine Learning, 2023. Corpus paper id: ICML 2023 0551; checked May 10, 2026
2023
-
[18]
Xiang Hu, Hongyu Fu, Jinge Wang, Yifeng Wang, Zhikun Li, Renjun Xu, Yu Lu, Yaochu Jin, Lili Pan, and Zhenzhong Lan. Nova: An iterative planning and search approach to enhance novelty and diversity of LLM generated ideas, 2024. arXiv:2410.14255; checked May 10, 2026
Pith/arXiv arXiv 2024
-
[19]
Jin Huang, Silviu Cucerzan, Sujay Kumar Jauhar, and Ryen W. White. Idea2plan: Exploring AI-powered research planning, 2025. arXiv:2510.24891; checked May 10, 2026
arXiv 2025
-
[20]
Hindsight: Evaluating LLM-generated research ideas via future impact, 2026
Bo Jiang. Hindsight: Evaluating LLM-generated research ideas via future impact, 2026. arXiv:2603.15164; checked May 10, 2026
arXiv 2026
-
[21]
Kingma and Ruiqi Gao
Diederik P. Kingma and Ruiqi Gao. Understanding diffusion objectives as the ELBO with simple data augmentation. InAdvances in Neural Information Processing Systems, 2023. Corpus paper id: NeurIPS 2023 0753; checked May 10, 2026. 50
2023
-
[22]
Fine-tuning can dis- tort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning can dis- tort pretrained features and underperform out-of-distribution. InInternational Conference on Learning Representations, 2022. Corpus paper id: ICLR 2022 0094; checked May 10, 2026
2022
-
[23]
Xinping Lei, Tong Zhou, Yubo Chen, Kang Liu, and Jun Zhao. Motivgraph-soiq: Integrating motivational knowledge graphs and socratic dialogue for enhanced LLM ideation, 2025. arXiv:2509.21978; checked May 10, 2026
arXiv 2025
-
[24]
Retrieval- augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Hein- rich Kuttler, Mike Lewis, Wen tau Yih, Tim Rocktaschel, Sebastian Riedel, and Douwe Kiela. Retrieval- augmented generation for knowledge-intensive NLP tasks. InAdvances in Neural Information Processing Systems, 2020. NeurIPS 2020; arXiv:2005.11401; checked...
Pith/arXiv arXiv 2020
-
[25]
MIRAI: Prediction and generation of high-impact academic research, 2026
Alex Li et al. MIRAI: Prediction and generation of high-impact academic research, 2026
2026
-
[26]
A vision for auto research with LLM agents, 2025
Chengwei Liu, Chong Wang, Jiayue Cao, Jingquan Ge, Kun Wang, Lyuye Zhang, Ming-Ming Cheng, Penghai Zhao, Tianlin Li, Xiaojun Jia, Xiang Li, Xingshuai Li, Yang Liu, Yebo Feng, Yihao Huang, Yijia Xu, Yuqiang Sun, Zhenhong Zhou, and Zhengzi Xu. A vision for auto research with LLM agents, 2025. arXiv:2504.18765; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[27]
Sci-reasoning: A dataset decoding AI innovation pat- terns, 2026
Jiachen Liu, Maestro Harmon, and Zechen Zhang. Sci-reasoning: A dataset decoding AI innovation pat- terns, 2026. arXiv:2601.04577; checked May 10, 2026
arXiv 2026
-
[28]
Augmenting research ideation with data: An empirical investigation in social science, 2025
Xiao Liu, Xinyi Dong, Xinyang Gao, Yansong Feng, and Xun Pang. Augmenting research ideation with data: An empirical investigation in social science, 2025. arXiv:2505.21396; checked May 10, 2026
arXiv 2025
-
[29]
Yujie Liu, Zonglin Yang, Tong Xie, Jinjie Ni, Ben Gao, Yuqiang Li, Shixiang Tang, Wanli Ouyang, Erik Cambria, and Dongzhan Zhou. Researchbench: Benchmarking LLMs in scientific discovery via inspiration- based task decomposition, 2025. arXiv:2503.21248; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[30]
Towards end-to-end automation of AI research.Nature, 2026
Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, and Jeff Clune. Towards end-to-end automation of AI research.Nature, 2026. Nature 2026; originally arXiv:2408.06292; checked May 10, 2026
Pith/arXiv arXiv 2026
-
[31]
hdbscan: Hierarchical density based clustering.Journal of Open Source Software, 2(11):205, 2017
Leland McInnes, John Healy, and Steve Astels. hdbscan: Hierarchical density based clustering.Journal of Open Source Software, 2(11):205, 2017. DOI:10.21105/joss.00205; checked May 10, 2026
-
[32]
UMAP: Uniform manifold approximation and projection for dimension reduction, 2018
Leland McInnes, John Healy, and James Melville. UMAP: Uniform manifold approximation and projection for dimension reduction, 2018. arXiv:1802.03426; checked May 10, 2026
Pith/arXiv arXiv 2018
-
[33]
SAM as an optimal relaxation of bayes
Thomas M¨ ollenhoff and Mohammad Emtiyaz Khan. SAM as an optimal relaxation of bayes. In International Conference on Learning Representations, 2023. ICLR 2023 Spotlight; corpus paper id: ICLR 2023 0331; checked May 10, 2026
2023
-
[34]
Embeddings, 2026
OpenAI. Embeddings, 2026. Official OpenAI API documentation; checked May 10, 2026
2026
-
[35]
Openreview documentation: Using the api, 2026
OpenReview. Openreview documentation: Using the api, 2026. Official OpenReview documentation; checked May 10, 2026
2026
-
[36]
AI idea bench 2025: AI research idea generation benchmark, 2025
Yansheng Qiu, Haoquan Zhang, Zhaopan Xu, Ming Li, Diping Song, Zheng Wang, and Kaipeng Zhang. AI idea bench 2025: AI research idea generation benchmark, 2025. arXiv:2504.14191; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[37]
Marissa Radensky, Simra Shahid, Raymond Fok, Pao Siangliulue, Tom Hope, and Daniel S. Weld. Human- LLM compound system for scientific ideation through facet recombination and novelty evaluation, 2024. Introduces Scideator; arXiv:2409.14634; checked May 10, 2026
Pith/arXiv arXiv 2024
-
[38]
Towards scientific intelli- gence: A survey of LLM-based scientific agents, 2025
Shuo Ren, Can Xie, Pu Jian, Zhenjiang Ren, Chunlin Leng, and Jiajun Zhang. Towards scientific intelli- gence: A survey of LLM-based scientific agents, 2025. arXiv:2503.24047; checked May 10, 2026
arXiv 2025
-
[39]
Evaluating novelty in AI-generated research plans using multi-workflow LLM pipelines, 2025
Devesh Saraogi, Rohit Singhee, and Dhruv Kumar. Evaluating novelty in AI-generated research plans using multi-workflow LLM pipelines, 2025. arXiv:2601.09714; checked May 10, 2026
arXiv 2025
-
[40]
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dess` ı, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, 2023. NeurIPS 2023; arXiv:2302.04761; checked May 10, 2026. 51
Pith/arXiv arXiv 2023
-
[41]
Agent laboratory: Using LLM agents as research assistants, 2025
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using LLM agents as research assistants, 2025. arXiv:2501.04227; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[42]
Is this idea novel? an automated benchmark for judgment of research ideas, 2026
Tim Schopf et al. Is this idea novel? an automated benchmark for judgment of research ideas, 2026
2026
-
[43]
Semantic scholar graph api, 2026
Semantic Scholar. Semantic scholar graph api, 2026. Official Semantic Scholar API documentation; checked May 10, 2026
2026
-
[44]
Simra Shahid, Marissa Radensky, Raymond Fok, Pao Siangliulue, Daniel S. Weld, and Tom Hope. Literature-grounded novelty assessment of scientific ideas, 2025. arXiv:2506.22026; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[45]
Yuexi Shen, Minqian Liu, Dawei Zhou, and Lifu Huang. Navigating ideation space: Decomposed conceptual representations for positioning scientific ideas, 2026. arXiv:2601.08901; checked May 10, 2026
arXiv 2026
-
[46]
The ideation-execution gap: Execution outcomes of LLM-generated versus human research ideas, 2025
Chenglei Si, Tatsunori Hashimoto, and Diyi Yang. The ideation-execution gap: Execution outcomes of LLM-generated versus human research ideas, 2025. arXiv:2506.20803; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[47]
Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers, 2024
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers, 2024. arXiv:2409.04109; checked May 10, 2026
Pith/arXiv arXiv 2024
-
[48]
Scirepeval: A multi- format benchmark for scientific document representations, 2022
Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman. Scirepeval: A multi- format benchmark for scientific document representations, 2022. arXiv:2211.13308; includes SPECTER2 evaluation context; checked May 10, 2026
Pith/arXiv arXiv 2022
-
[49]
Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, and Nanqing Dong. Many heads are better than one: Improved scientific idea generation by a LLM-based multi-agent system, 2024. arXiv:2410.09403; checked May 10, 2026
Pith/arXiv arXiv 2024
-
[50]
AI-researcher: Autonomous scientific inno- vation, 2025
Jiabin Tang, Lianghao Xia, Zhonghang Li, and Chao Huang. AI-researcher: Autonomous scientific inno- vation, 2025. arXiv:2505.18705; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[51]
AI can learn scientific taste, 2026
Jingqi Tong, Mingzhe Li, Hangcheng Li, Yongzhuo Yang, Yurong Mou, Weijie Ma, Zhiheng Xi, Hongji Chen, Xiaoran Liu, Qinyuan Cheng, Ming Zhang, Qiguang Chen, Weifeng Ge, Qipeng Guo, Tianlei Ying, Tianxiang Sun, Yining Zheng, Xinchi Chen, Jun Zhao, Ning Ding, Xuanjing Huang, Yugang Jiang, and Xipeng Qiu. AI can learn scientific taste, 2026. arXiv:2603.14473;...
arXiv 2026
-
[52]
Why LLMs aren’t scientists yet: Lessons from four autonomous research attempts, 2026
Dhruv Trehan et al. Why LLMs aren’t scientists yet: Lessons from four autonomous research attempts, 2026
2026
-
[53]
Learning to predict future-aligned research proposals with language models, 2026
Heng Wang, Pengcheng Jiang, Jiashuo Sun, Zhiyi Shi, Haofei Yu, Jiawei Han, and Heng Ji. Learning to predict future-aligned research proposals with language models, 2026. arXiv:2603.27146; checked May 10, 2026
Pith/arXiv arXiv 2026
-
[54]
FlowPIE: Test-time scientific idea evolution with flow-guided literature exploration, 2026
Qiyao Wang et al. FlowPIE: Test-time scientific idea evolution with flow-guided literature exploration, 2026
2026
-
[55]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. InAdvances in Neural Information Processing Systems, 2022. NeurIPS 2022; arXiv:2201.11903; checked May 10, 2026
Pith/arXiv arXiv 2022
-
[56]
Novbench: Evaluating large language models on academic paper novelty assessment, 2026
Wenqing Wu, Yi Zhao, Yuzhuo Wang, Siyou Li, Juexi Shao, Yunfei Long, and Chengzhi Zhang. Novbench: Evaluating large language models on academic paper novelty assessment, 2026. arXiv:2604.11543; checked May 10, 2026
Pith/arXiv arXiv 2026
-
[57]
Improving scientific hypothesis generation with knowledge grounded large language models, 2024
Guangzhi Xiong, Eric Xie, Amir Hassan Shariatmadari, Sikun Guo, Stefan Bekiranov, and Aidong Zhang. Improving scientific hypothesis generation with knowledge grounded large language models, 2024. arXiv:2411.02382; checked May 10, 2026
Pith/arXiv arXiv 2024
-
[58]
The AI scientist-v2: Workshop-level automated scientific discovery via agentic tree search, 2025
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The AI scientist-v2: Workshop-level automated scientific discovery via agentic tree search, 2025. arXiv:2504.08066; checked May 10, 2026
Pith/arXiv arXiv 2025
-
[59]
ReAct: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. InInternational Conference on Learning Represen- tations, 2023. ICLR 2023; arXiv:2210.03629; checked May 10, 2026. 52
Pith/arXiv arXiv 2023
-
[60]
Opennovelty: An LLM-powered agentic system for verifiable scholarly novelty assessment, 2026
Ming Zhang, Kexin Tan, Yueyuan Huang, Yujiong Shen, Chunchun Ma, Li Ju, Xinran Zhang, Yuhui Wang, Wenqing Jing, Jingyi Deng, Huayu Sha, Binze Hu, Jingqi Tong, Changhao Jiang, Yage Geng, Yuankai Ying, Yue Zhang, Zhangyue Yin, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, and Xuanjing Huang. Opennovelty: An LLM-powered agentic system for verifiable scholarly n...
arXiv 2026
-
[61]
Keyu Zhao, Weiquan Lin, Qirui Zheng, Fengli Xu, and Yong Li. Deep ideation: Designing LLM agents to generate novel research ideas on scientific concept network, 2025. arXiv:2511.02238; checked May 10, 2026
arXiv 2025
-
[62]
From automation to autonomy: A survey on large language models in scientific discovery, 2025
Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song. From automation to autonomy: A survey on large language models in scientific discovery, 2025. arXiv:2505.13259; checked May 10, 2026
arXiv 2025
-
[63]
Tianshi Zheng, Kelvin Kiu-Wai Tam, Newt Hue-Nam K. Nguyen, Baixuan Xu, Zhaowei Wang, Jiayang Cheng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, and Simon See. Newtonbench: Benchmarking generalizable scientific law discovery in LLM agents, 2025. arXiv:2510.07172; checked May 10, 2026. 53
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.