Pith. sign in

REVIEW 3 major objections 6 minor 63 references

Public ML conference outcomes contain reusable ideation patterns that can be turned into skills for writing stronger, evidence-grounded research proposals.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 19:10 UTC pith:6MIIPCJT

load-bearing objection Solid systems paper: contrastive Oral/HC/Reject pattern cards plus a real first-mile skill suite; the quality win is real inside its automated protocol but still partly style-aligned with the judge. the 3 major comments →

arxiv 2607.04439 v1 pith:6MIIPCJT submitted 2026-07-05 cs.AI

ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes

classification cs.AI
keywords research ideationLLM agentsideation patternsprior-art collisionconference outcomesskill suitenovelty evaluationmachine learning research
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that the first mile of research—grounding a problem in literature, naming a real bottleneck, differentiating from prior work, and auditing risk—can be organized as reusable skills rather than free-form brainstorming. From 1,947 ICLR, ICML, and NeurIPS papers (2021–2025), including Oral, high-citation, and rejected work, the authors induce 31 fine-grained moves and consolidate them into 15 ideation patterns. Each pattern is packaged as a card that records when it applies, how successful papers execute it, how rejected papers fail it, and what reviewers typically demand. IdeaSpark composes multi-source literature grounding, bottleneck diagnosis, pattern-guided generation, prior-art collision checking, and outcome-informed audit into one workflow that outputs a single auditable idea card. In blind automated-judge tests on 100 held-out ICLR 2026 Oral seeds, IdeaSpark ranks highest on idea quality against bare-model and generic-skill baselines while remaining competitively novel, supporting the claim that conference outcomes carry operational signals about how impactful directions are formulated and checked.

Core claim

The authors claim that Oral, high-citation, and rejected ML conference papers share a compact strategy space of about 15 reusable ideation patterns, that rejected work is best read as contrastive evidence about weak instantiations rather than a separate strategy class, and that packaging those patterns as structured cards inside a multi-phase, retrieval-and-audit skill produces research proposals that automated judges rate higher in quality than no-skill or generic-skill baselines while remaining competitively novel.

What carries the argument

The ideation-pattern card: a structured, corpus-derived object that pairs operational signature, when-to-apply conditions, success conditions from accepted work, failure modes from rejected work, and reviewer expectations—used both to select a research move for a diagnosed bottleneck and to audit the generated candidate against the same corpus-derived failure modes.

Load-bearing premise

That two automated LLM judging skills for idea quality and prior-art collision are close enough to human expert judgment to support the endpoint quality claim.

What would settle it

A blind human-expert ranking study on the same 100 method-agnostic ICLR 2026 seeds in which IdeaSpark no longer wins on quality relative to the same baselines, or in which human novelty judgments reverse the competitive-novelty result.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Early-stage research ideation can be treated as an evidence-grounded skill layer rather than open-ended brainstorming.
  • Accepted and rejected ML papers often use the same high-level moves; the useful signal is execution quality and failure modes, not strategy choice alone.
  • Pattern composition of one to three moves, with two as the modal default, better matches real papers than single-move recipes.
  • Standalone literature-search and prior-art-collision skills can be reused outside the full ideation loop.
  • Automated quality–novelty planes can expose “novel-but-empty” proposals that score high on novelty only because they are too vague to collide with prior art.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same accept/reject contrastive-card idea could transfer to other fields that publish decisions and reviews, not only mainstream ML conferences.
  • If human experts confirm the automated quality ranking, labs could use pattern cards as pre-experiment checklists rather than only as generation prompts.
  • The finding that rejected constructive moves often drift toward audit/unify surface forms suggests a practical diagnostic for early drafts: check whether a construction has collapsed into a weaker audit claim.
  • Separating generation freedom from audit anchoring may be a general design rule for other LLM research tools that currently over-constrain generation with corpus frequency priors.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces ResearchStudio-Idea, a three-skill suite (Paper-Search, Scoop-Check, IdeaSpark) for the first mile of ML research ideation. From 1,947 ICLR/ICML/NeurIPS papers (2021–2025) labeled Oral, high-citation, and Reject, the authors extract domain-agnostic innovation signatures, cluster them into 31 sub-patterns, and induce 15 operational ideation-pattern cards that encode success conditions, failure modes, and differentiation strategies. IdeaSpark composes literature grounding, bottleneck diagnosis, pattern-guided candidate construction, prior-art collision retrieval, and outcome-informed audit into a single idea-card workflow with deterministic validators. A blind automated-judge study on 100 method-agnostic ICLR 2026 Oral seeds reports that IdeaSpark attains mean quality 3.87/4 (first on 88/100 seeds) versus same-backbone bare and generic-skill baselines, while remaining competitively novel on a scoop-check scale.

Significance. If the results hold, the work supplies a concrete missing layer between conference outcome data and inference-time ideation: reusable, failure-aware pattern cards induced from Oral/HC/Reject contrast rather than accepted-only taxonomies or parametric policies. Strengths include the three-way outcome contrast, reject-only remapping showing shared strategy space (§10), multi-label composition analysis with k=2 as mode (§6.5), embedding/abstraction ablations (§11), same-backbone evaluation ladder with length/format normalization (§14), and explicit faithfulness machinery (kill-switch fields, citation gates, honest abandon paths). The quality–novelty plane analysis that exposes a “novel-but-empty” failure mode is a useful methodological contribution beyond the system itself. The main open question is whether automated-judge gains transfer to human expert preference and downstream research value.

major comments (3)
  1. [§14 Evaluation] §14.4–§14.7 and Tables 11–12: The endpoint quality claim rests on an idea-quality judge that rewards problem position, method depth (dominating), soundness/feasibility, and problem-fit, with formal equations required by the normalization contract (§14.3). IdeaSpark is explicitly engineered to emit that surface (bottleneck diagnosis, sub-pattern instantiation, equations with interpretation, falsification/kill-switch, implementability audit; §13.3–13.4, App. A). The same-backbone ladder isolates skill design for generation, but both judges are LLM skills and no control rewrites IdeaSpark cards into bare-style prose (or expands bare outputs into IdeaSpark structure) before re-ranking. Without that control, cross-family judges, or a human pilot, the 3.87/4 and 88/100 wins may partly measure preference for the skill’s own output contract rather than independent research-move quality—the feasi
  2. [§6 Ideation-Pattern Induction] §1.4 and §6: The 15-pattern taxonomy is produced by a single Opus 4.7 induction call with a 6–18 size band; the paper notes that a comparable run would likely land at 12–18 with substantial overlap, but reports no inter-prompt, inter-seed, or inter-model stability study. Because the cards are the shared object for both generation and audit, instability in pattern boundaries would affect both Phase 2 selection and Phase 3 failure-mode checks. A modest multi-seed or multi-model remapping (e.g., Jaccard of cluster→pattern assignments, or card-level content agreement on a held-out paper sample) is needed to support the claim that the induced map is a reusable operational substrate rather than a single-run artifact.
  3. [§14.2 Problem seeds / §15 Limitations] §14.2 and §15: Evaluation seeds are method-agnostic rewrites of ICLR 2026 Oral titles. While held-out relative to the 2021–2025 induction corpus, they still sample the same venue’s Oral distribution and primary-area mix (Fig. 1 right). Combined with corpus restriction to three ML conferences, this leaves open whether quality gains hold for non-Oral problem framings, systems/benchmark-style directions that the quality rubric systematically down-weights (§15), or domains outside the induced 28-area map. The paper should either (i) add a non-Oral or cross-venue seed set, or (ii) more tightly scope the endpoint claim to “Oral-like ML conference directions under automated judges.”
minor comments (6)
  1. [Abstract / §1.4] Abstract and §1.4: Soften absolute phrasing (“consistently produces stronger research proposals”) to match the paper’s own scoping as an automated-judge, idea-stage study until human or style-control evidence is added.
  2. [§5.2 Clustering] §5.2 Table 2: 47.7% unclustered at the selected min_cluster_size is high; the multi-label pass (§6.5) addresses coverage, but a short quantitative comparison of multi-label pattern distributions for clustered vs unclustered papers would reassure readers that unclustered points are not systematically different strategies.
  3. [§7 Acceptance and Impact Analysis] §7 Table 7: ∆_OR uses cluster-level primaries only (N_O=482, N_R=367). State more prominently in the table caption that these are not full-class denominators, to avoid over-reading the tight ±2.9 pp spread.
  4. [Figure 1] Figure 1 left: Axis labels and baseline names are dense; a legend mapping “Opus-4.8 (self-gen)” to the generic auto-authored skill would help readers who land on the figure first.
  5. [§3.1 Scope and labeling] §3.1: Clarify how the 49 Oral∩HC overlaps are handled in every analysis that reports mutually exclusive classes versus inclusive HC_all, ideally with a single convention table.
  6. [Appendix] App. A–D are valuable; ensure the main text points to which appendix artifact corresponds to which evaluation seed ID so readers can audit a full run trail.

Circularity Check

0 steps flagged

No load-bearing circular derivation; mild residual that automated quality axes align with IdeaSpark’s engineered card contract, which the paper itself scopes as an automated-judge study.

full rationale

This paper does not claim a first-principles mathematical prediction that reduces to its inputs by construction. The empirical chain is corpus construction → strategy-signature extraction → clustering → 15/31 pattern cards → inference-time skill, then a separate held-out endpoint evaluation on method-agnostic ICLR 2026 Oral seeds. Evaluation seeds post-date the 2021–2025 induction corpus and strip each paper’s own method, so the quality/novelty results are not rediscovery of the training set. Pattern cards dual-used for generation and audit are an intentional design choice, not a self-definitional proof. Scoop-Check appears both as a suite skill and as the novelty judge, but novelty scores are not fitted targets that force IdeaSpark’s generation, and IdeaSpark is not the novelty leader (GPT-5.5 bare is). Format normalization (matched Title/Motivation/Method lengths; required formal equations for all sources) and the same-backbone ladder (Opus bare ≈ Opus-self-gen ≪ IdeaSpark) further block a pure “format equals quality” reduction. The residual concern—that idea-quality’s depth/problem-fit axes philosophically match what IdeaSpark is built to emit—is a construct-validity / human-alignment limitation the paper already states in §14–§15, not a circular step of the kinds this pass flags. No self-citation uniqueness theorem, fitted-as-prediction, or renamed known law anchors the central claim.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

This is an empirical systems paper, not a formal derivation. The central claim rests on public conference outcomes as traces of research practice, on LLM extraction/induction as faithful strategy distillers, on clustering geometry as a proxy for strategy similarity, and on automated judges as proxies for idea quality/novelty. Free parameters are mostly pipeline hyperparameters and model choices rather than fitted physical constants. The main invented entities are the operational pattern cards and skill phases themselves.

free parameters (5)
  • HDBSCAN min_cluster_size
    Chosen as 10 after a silhouette sweep; directly determines the 31-cluster inventory used for pattern induction.
  • UMAP n_neighbors / n_components / min_dist
    Fixed embedding-reduction settings that shape cluster geometry before HDBSCAN.
  • Target taxonomy size band (6–18 patterns)
    Prompt constraint on Opus induction; the returned 15-pattern cut is a single-call operational granularity, not a unique optimum.
  • Oral-safe / Reject-warn thresholds (p_O ≥65% / ≤35%)
    Hand-set risk flags used in descriptive cluster analysis and audit context.
  • Idea-quality and scoop-check scoring rubrics
    Judge-axis definitions and worst-case novelty aggregation determine the endpoint quality/novelty numbers.
axioms (5)
  • domain assumption Public ML conference outcomes (Oral, high-citation, Reject) contain reusable traces of how impactful research directions are formulated, differentiated, and evaluated.
    Stated as the empirical premise of the whole project in the introduction and conclusion.
  • domain assumption Domain-agnostic rewriting of innovation fields yields strategy clusters rather than topic clusters.
    Justified by the Stage-2 rewrite design and the embedding ablation in §11.
  • domain assumption LLM extraction, multi-label tagging, and pattern-card induction preserve operationally useful strategy structure.
    All signature fields, taxonomy induction, and cards are model-generated under schemas; no independent human annotation gold set is reported.
  • domain assumption Automated idea-quality and scoop-check judges provide a useful first-stage ranking of proposal quality and novelty.
    Endpoint evaluation in §14 depends entirely on these judges; human agreement is deferred.
  • standard math Standard embedding/clustering mathematics (cosine similarity, UMAP, HDBSCAN, silhouette) is a valid substrate for grouping research strategies.
    Used as off-the-shelf tools without novel theorems.
invented entities (3)
  • 15 ideation-pattern cards / 31 sub-pattern cards no independent evidence
    purpose: Serve as the shared non-parametric object for pattern selection, candidate instantiation, and failure-mode audit.
    Induced constructs, not pre-existing scientific objects; independent_evidence is only internal corpus coherence and automated endpoint gains.
  • IdeaSpark multi-phase skill (Phases 0–4 with kill-switch validators) no independent evidence
    purpose: Turn pattern cards into a runnable research-problem-to-idea-card workflow with retrieval gates and deterministic checks.
    System artifact introduced by the paper; evidence is the described implementation and automated evaluation, not external adoption yet.
  • Scoop-Check four-axis novelty collision level no independent evidence
    purpose: Score prior-art overlap on problem framing, core mechanism, key insight, and application domain.
    Evaluation instrument defined for this study; related novelty tools exist, but this specific skill/scale is paper-local.

pith-pipeline@v1.1.0-grok45 · 54682 in / 3749 out tokens · 50306 ms · 2026-07-11T19:10:05.941918+00:00 · methodology

0 comments
read the original abstract

Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating candidate directions. Researchers must ground a problem in current literature, identify meaningful bottlenecks, differentiate from existing solutions, and evaluate risks before committing to implementation. We present ResearchStudio-Idea as a reusable skill suite for this first mile of research ideation. The suite includes Paper-Search, a standalone multi-source literature search skill; Scoop-Check, a standalone prior-art collision checker for novelty claims; and IdeaSpark, the end-to-end skill that composes evidence grounding, pattern-guided generation, collision retrieval, audit, and idea-card rendering into one workflow. IdeaSpark is constructed from a corpus of 1,947 machine learning conference papers collected from ICLR, ICML, and NeurIPS between 2021 and 2025, including Oral papers, a separately tracked high-citation subset, and rejected submissions. Analysis of these outcomes reveals 31 recurring ideation sub-patterns, consolidated into 15 reusable ideation patterns. Each pattern is operationalized as a structured card containing research contexts, bottleneck types, differentiation strategies, supporting precedents, and common failure modes. Given a research problem and an evidence bundle, IdeaSpark evaluates evidence readiness, reconstructs the surrounding research context, identifies unresolved bottlenecks, selects relevant patterns, instantiates one candidate direction, retrieves potentially conflicting prior work, and performs outcome-informed auditing. This workflow transforms reusable ideation patterns into traceable research proposals. Blind automated-judge evaluations show that IdeaSpark consistently produces stronger research proposals than no-skill and generic-skill baselines while maintaining competitive novelty.

Figures

Figures reproduced from arXiv: 2607.04439 by Jianjun Gao, Lingao Xiao, Qihao Zhao, Scarlett Li, Wenshan Wu, Xin Zhang, Yalun Dai, Yang He, Yangyu Huang, Yan Lu, Yap Kim Hui.

Figure 1
Figure 1. Figure 1: IdeaSpark improves idea quality while maintaining competitive novelty in blind automated-judge evaluations. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: IdeaSpark data-to-skill workflow. The upper band constructs reusable ideation assets from the [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Acceptance composition of the 31 clusters, sorted by Oral rate among [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: 2D UMAP projection of the 1,891-paper embedding. Oral, HC, and Reject papers are interleaved [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Ideation-pattern hierarchy across 989 clustered papers. Three concentric rings share one angular layout, ordered by size clockwise from 12 o’clock; categorical colors identify the 15 patterns and are reused across all rings, with white gaps separating segments. Inner ring: the 15 induced Level-1 ideation patterns, each wedge sized by its clustered-paper count (the count is printed inside the larger wedges;… view at source ↗
Figure 6
Figure 6. Figure 6: Centroid-cosine similarity between the 15 ideation patterns. Rows/columns ordered by mean similarity [PITH_FULL_IMAGE:figures/full_fig_p018_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Multi-label ideation pattern assignment. (a) [PITH_FULL_IMAGE:figures/full_fig_p019_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Top combinations under multi-label tagging. (a) 2-way combinations ( [PITH_FULL_IMAGE:figures/full_fig_p020_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Per-ideation-pattern acceptance bias. Bars show ∆ [PITH_FULL_IMAGE:figures/full_fig_p022_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: (a) Paper counts per (domain, ideation pattern) cell, with domains and ideation patterns sorted by [PITH_FULL_IMAGE:figures/full_fig_p023_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Ideation-pattern breadth: number of distinct domains each ideation pattern has [PITH_FULL_IMAGE:figures/full_fig_p024_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Per-pattern share of papers by year under multi-label tagging (% of that year’s papers carrying the [PITH_FULL_IMAGE:figures/full_fig_p025_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Pattern usage by venue under multi-label tagging. Each bar reports the share of papers in a venue [PITH_FULL_IMAGE:figures/full_fig_p026_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Left: Reject-only clustering (n=711, 13 clusters, 40.5% unclustered). Right: Reject-only-clustering → main mapping vs. main taxonomy primary distribution over the same Reject papers. The two columns largely agree; the divergences are concentrated on Unify Heterogeneous Inputs into One Space and Audit and Pivot an Assumption (over-attracted) vs. Encode Structure by Construction (absent from the Reject-only… view at source ↗
Figure 15
Figure 15. Figure 15: Clustering quality across three embedding-input choices, as min [PITH_FULL_IMAGE:figures/full_fig_p029_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Distribution of novelty levels per system (L1 = fully scooped [PITH_FULL_IMAGE:figures/full_fig_p038_16.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

63 extracted references · 30 linked inside Pith

  1. [1]

    SPECTER2 model card, 2026

    Allen Institute for AI. SPECTER2 model card, 2026. Official SPECTER2 Hugging Face model card; checked May 10, 2026

  2. [2]

    Claude models overview, 2026

    Anthropic. Claude models overview, 2026. Official Anthropic model documentation; checked May 10, 2026

  3. [3]

    Skill authoring best practices, 2026

    Anthropic. Skill authoring best practices, 2026. Official Anthropic documentation; checked May 10, 2026

  4. [4]

    Artiles, Martin Weiss, Levin Brinkmann, Anirudh Goyal, and Nasim Rahaman

    Alejandro H. Artiles, Martin Weiss, Levin Brinkmann, Anirudh Goyal, and Nasim Rahaman. Alien science: Sampling coherent but cognitively unavailable research directions from idea atoms, 2026. arXiv:2603.01092; checked May 10, 2026

  5. [5]

    Agentic AI scientists are not built for autonomous scientific discovery, 2026

    Harshit Bisht et al. Agentic AI scientists are not built for autonomous scientific discovery, 2026

  6. [6]

    Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. SPECTER: Document-level representation learning using citation-informed transformers. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2270–2282, 2020. ACL 2020; arXiv:2004.07180; checked May 10, 2026

  7. [7]

    Graphmind: Interactive novelty assessment system for accelerating scientific discovery, 2025

    Italo Luis da Silva, Hanqi Yan, Lin Gui, and Yulan He. Graphmind: Interactive novelty assessment system for accelerating scientific discovery, 2025. arXiv:2510.15706; checked May 10, 2026

  8. [8]

    Graphs of research: Citation evolution graphs as supervision for research idea gener- ation, 2026

    Songyang Gao et al. Graphs of research: Citation evolution graphs as supervision for research idea gener- ation, 2026

  9. [9]

    Iris: Interactive research ideation system for accelerating scientific discovery, 2025

    Aniketh Garikaparthi, Manasi Patwardhan, Lovekesh Vig, and Arman Cohan. Iris: Interactive research ideation system for accelerating scientific discovery, 2025. arXiv:2504.16728; checked May 10, 2026

  10. [10]

    EigenGame: PCA as a nash equilibrium

    Ian Gemp, Brian McWilliams, Claire Vernade, and Thore Graepel. EigenGame: PCA as a nash equilibrium. InInternational Conference on Learning Representations, 2021. ICLR 2021 Outstanding Paper; corpus paper id: ICLR 2021 0018; checked May 10, 2026

  11. [11]

    Towards an AI co-scientist,

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, et al. Towards an AI co-scientist,

  12. [12]

    arXiv:2502.18864; checked May 10, 2026

  13. [13]

    BERTopic: Neural topic modeling with a class-based tf-idf procedure, 2022

    Maarten Grootendorst. BERTopic: Neural topic modeling with a class-based tf-idf procedure, 2022. arXiv:2203.05794; checked May 10, 2026

  14. [14]

    Mori: Learn- ing motivation-grounded reasoning for scientific ideation in large language models, 2026

    Chenyang Gu, Jiahao Cheng, Meicong Zhang, Pujun Zheng, Jinquan Zheng, and Guoxiu He. Mori: Learn- ing motivation-grounded reasoning for scientific ideation in large language models, 2026. arXiv:2603.19044; checked May 10, 2026

  15. [15]

    Interesting scientific idea generation using knowledge graphs and LLMs: Evaluations with 100 research group leaders, 2024

    Xuemei Gu and Mario Krenn. Interesting scientific idea generation using knowledge graphs and LLMs: Evaluations with 100 research group leaders, 2024. arXiv:2405.17044; checked May 10, 2026

  16. [16]

    Ideabench: Benchmarking large language models for research idea generation, 2024

    Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, Eric Xie, Stefan Bekiranov, and Aidong Zhang. Ideabench: Benchmarking large language models for research idea generation, 2024. arXiv:2411.02429; checked May 10, 2026

  17. [17]

    The price of differential privacy under continual observation

    Monika Henzinger, Satish Sricharan, and Teresa Steiger. The price of differential privacy under continual observation. InInternational Conference on Machine Learning, 2023. Corpus paper id: ICML 2023 0551; checked May 10, 2026

  18. [18]

    Nova: An iterative planning and search approach to enhance novelty and diversity of LLM generated ideas, 2024

    Xiang Hu, Hongyu Fu, Jinge Wang, Yifeng Wang, Zhikun Li, Renjun Xu, Yu Lu, Yaochu Jin, Lili Pan, and Zhenzhong Lan. Nova: An iterative planning and search approach to enhance novelty and diversity of LLM generated ideas, 2024. arXiv:2410.14255; checked May 10, 2026

  19. [19]

    Jin Huang, Silviu Cucerzan, Sujay Kumar Jauhar, and Ryen W. White. Idea2plan: Exploring AI-powered research planning, 2025. arXiv:2510.24891; checked May 10, 2026

  20. [20]

    Hindsight: Evaluating LLM-generated research ideas via future impact, 2026

    Bo Jiang. Hindsight: Evaluating LLM-generated research ideas via future impact, 2026. arXiv:2603.15164; checked May 10, 2026

  21. [21]

    Kingma and Ruiqi Gao

    Diederik P. Kingma and Ruiqi Gao. Understanding diffusion objectives as the ELBO with simple data augmentation. InAdvances in Neural Information Processing Systems, 2023. Corpus paper id: NeurIPS 2023 0753; checked May 10, 2026. 50

  22. [22]

    Fine-tuning can dis- tort pretrained features and underperform out-of-distribution

    Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning can dis- tort pretrained features and underperform out-of-distribution. InInternational Conference on Learning Representations, 2022. Corpus paper id: ICLR 2022 0094; checked May 10, 2026

  23. [23]

    Motivgraph-soiq: Integrating motivational knowledge graphs and socratic dialogue for enhanced LLM ideation, 2025

    Xinping Lei, Tong Zhou, Yubo Chen, Kang Liu, and Jun Zhao. Motivgraph-soiq: Integrating motivational knowledge graphs and socratic dialogue for enhanced LLM ideation, 2025. arXiv:2509.21978; checked May 10, 2026

  24. [24]

    Retrieval- augmented generation for knowledge-intensive NLP tasks

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Hein- rich Kuttler, Mike Lewis, Wen tau Yih, Tim Rocktaschel, Sebastian Riedel, and Douwe Kiela. Retrieval- augmented generation for knowledge-intensive NLP tasks. InAdvances in Neural Information Processing Systems, 2020. NeurIPS 2020; arXiv:2005.11401; checked...

  25. [25]

    MIRAI: Prediction and generation of high-impact academic research, 2026

    Alex Li et al. MIRAI: Prediction and generation of high-impact academic research, 2026

  26. [26]

    A vision for auto research with LLM agents, 2025

    Chengwei Liu, Chong Wang, Jiayue Cao, Jingquan Ge, Kun Wang, Lyuye Zhang, Ming-Ming Cheng, Penghai Zhao, Tianlin Li, Xiaojun Jia, Xiang Li, Xingshuai Li, Yang Liu, Yebo Feng, Yihao Huang, Yijia Xu, Yuqiang Sun, Zhenhong Zhou, and Zhengzi Xu. A vision for auto research with LLM agents, 2025. arXiv:2504.18765; checked May 10, 2026

  27. [27]

    Sci-reasoning: A dataset decoding AI innovation pat- terns, 2026

    Jiachen Liu, Maestro Harmon, and Zechen Zhang. Sci-reasoning: A dataset decoding AI innovation pat- terns, 2026. arXiv:2601.04577; checked May 10, 2026

  28. [28]

    Augmenting research ideation with data: An empirical investigation in social science, 2025

    Xiao Liu, Xinyi Dong, Xinyang Gao, Yansong Feng, and Xun Pang. Augmenting research ideation with data: An empirical investigation in social science, 2025. arXiv:2505.21396; checked May 10, 2026

  29. [29]

    Researchbench: Benchmarking LLMs in scientific discovery via inspiration- based task decomposition, 2025

    Yujie Liu, Zonglin Yang, Tong Xie, Jinjie Ni, Ben Gao, Yuqiang Li, Shixiang Tang, Wanli Ouyang, Erik Cambria, and Dongzhan Zhou. Researchbench: Benchmarking LLMs in scientific discovery via inspiration- based task decomposition, 2025. arXiv:2503.21248; checked May 10, 2026

  30. [30]

    Towards end-to-end automation of AI research.Nature, 2026

    Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, and Jeff Clune. Towards end-to-end automation of AI research.Nature, 2026. Nature 2026; originally arXiv:2408.06292; checked May 10, 2026

  31. [31]

    hdbscan: Hierarchical density based clustering.Journal of Open Source Software, 2(11):205, 2017

    Leland McInnes, John Healy, and Steve Astels. hdbscan: Hierarchical density based clustering.Journal of Open Source Software, 2(11):205, 2017. DOI:10.21105/joss.00205; checked May 10, 2026

  32. [32]

    UMAP: Uniform manifold approximation and projection for dimension reduction, 2018

    Leland McInnes, John Healy, and James Melville. UMAP: Uniform manifold approximation and projection for dimension reduction, 2018. arXiv:1802.03426; checked May 10, 2026

  33. [33]

    SAM as an optimal relaxation of bayes

    Thomas M¨ ollenhoff and Mohammad Emtiyaz Khan. SAM as an optimal relaxation of bayes. In International Conference on Learning Representations, 2023. ICLR 2023 Spotlight; corpus paper id: ICLR 2023 0331; checked May 10, 2026

  34. [34]

    Embeddings, 2026

    OpenAI. Embeddings, 2026. Official OpenAI API documentation; checked May 10, 2026

  35. [35]

    Openreview documentation: Using the api, 2026

    OpenReview. Openreview documentation: Using the api, 2026. Official OpenReview documentation; checked May 10, 2026

  36. [36]

    AI idea bench 2025: AI research idea generation benchmark, 2025

    Yansheng Qiu, Haoquan Zhang, Zhaopan Xu, Ming Li, Diping Song, Zheng Wang, and Kaipeng Zhang. AI idea bench 2025: AI research idea generation benchmark, 2025. arXiv:2504.14191; checked May 10, 2026

  37. [37]

    Marissa Radensky, Simra Shahid, Raymond Fok, Pao Siangliulue, Tom Hope, and Daniel S. Weld. Human- LLM compound system for scientific ideation through facet recombination and novelty evaluation, 2024. Introduces Scideator; arXiv:2409.14634; checked May 10, 2026

  38. [38]

    Towards scientific intelli- gence: A survey of LLM-based scientific agents, 2025

    Shuo Ren, Can Xie, Pu Jian, Zhenjiang Ren, Chunlin Leng, and Jiajun Zhang. Towards scientific intelli- gence: A survey of LLM-based scientific agents, 2025. arXiv:2503.24047; checked May 10, 2026

  39. [39]

    Evaluating novelty in AI-generated research plans using multi-workflow LLM pipelines, 2025

    Devesh Saraogi, Rohit Singhee, and Dhruv Kumar. Evaluating novelty in AI-generated research plans using multi-workflow LLM pipelines, 2025. arXiv:2601.09714; checked May 10, 2026

  40. [40]

    Toolformer: Language models can teach themselves to use tools

    Timo Schick, Jane Dwivedi-Yu, Roberto Dess` ı, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, 2023. NeurIPS 2023; arXiv:2302.04761; checked May 10, 2026. 51

  41. [41]

    Agent laboratory: Using LLM agents as research assistants, 2025

    Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using LLM agents as research assistants, 2025. arXiv:2501.04227; checked May 10, 2026

  42. [42]

    Is this idea novel? an automated benchmark for judgment of research ideas, 2026

    Tim Schopf et al. Is this idea novel? an automated benchmark for judgment of research ideas, 2026

  43. [43]

    Semantic scholar graph api, 2026

    Semantic Scholar. Semantic scholar graph api, 2026. Official Semantic Scholar API documentation; checked May 10, 2026

  44. [44]

    Weld, and Tom Hope

    Simra Shahid, Marissa Radensky, Raymond Fok, Pao Siangliulue, Daniel S. Weld, and Tom Hope. Literature-grounded novelty assessment of scientific ideas, 2025. arXiv:2506.22026; checked May 10, 2026

  45. [45]

    Navigating ideation space: Decomposed conceptual representations for positioning scientific ideas, 2026

    Yuexi Shen, Minqian Liu, Dawei Zhou, and Lifu Huang. Navigating ideation space: Decomposed conceptual representations for positioning scientific ideas, 2026. arXiv:2601.08901; checked May 10, 2026

  46. [46]

    The ideation-execution gap: Execution outcomes of LLM-generated versus human research ideas, 2025

    Chenglei Si, Tatsunori Hashimoto, and Diyi Yang. The ideation-execution gap: Execution outcomes of LLM-generated versus human research ideas, 2025. arXiv:2506.20803; checked May 10, 2026

  47. [47]

    Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers, 2024

    Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers, 2024. arXiv:2409.04109; checked May 10, 2026

  48. [48]

    Scirepeval: A multi- format benchmark for scientific document representations, 2022

    Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman. Scirepeval: A multi- format benchmark for scientific document representations, 2022. arXiv:2211.13308; includes SPECTER2 evaluation context; checked May 10, 2026

  49. [49]

    Many heads are better than one: Improved scientific idea generation by a LLM-based multi-agent system, 2024

    Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, and Nanqing Dong. Many heads are better than one: Improved scientific idea generation by a LLM-based multi-agent system, 2024. arXiv:2410.09403; checked May 10, 2026

  50. [50]

    AI-researcher: Autonomous scientific inno- vation, 2025

    Jiabin Tang, Lianghao Xia, Zhonghang Li, and Chao Huang. AI-researcher: Autonomous scientific inno- vation, 2025. arXiv:2505.18705; checked May 10, 2026

  51. [51]

    AI can learn scientific taste, 2026

    Jingqi Tong, Mingzhe Li, Hangcheng Li, Yongzhuo Yang, Yurong Mou, Weijie Ma, Zhiheng Xi, Hongji Chen, Xiaoran Liu, Qinyuan Cheng, Ming Zhang, Qiguang Chen, Weifeng Ge, Qipeng Guo, Tianlei Ying, Tianxiang Sun, Yining Zheng, Xinchi Chen, Jun Zhao, Ning Ding, Xuanjing Huang, Yugang Jiang, and Xipeng Qiu. AI can learn scientific taste, 2026. arXiv:2603.14473;...

  52. [52]

    Why LLMs aren’t scientists yet: Lessons from four autonomous research attempts, 2026

    Dhruv Trehan et al. Why LLMs aren’t scientists yet: Lessons from four autonomous research attempts, 2026

  53. [53]

    Learning to predict future-aligned research proposals with language models, 2026

    Heng Wang, Pengcheng Jiang, Jiashuo Sun, Zhiyi Shi, Haofei Yu, Jiawei Han, and Heng Ji. Learning to predict future-aligned research proposals with language models, 2026. arXiv:2603.27146; checked May 10, 2026

  54. [54]

    FlowPIE: Test-time scientific idea evolution with flow-guided literature exploration, 2026

    Qiyao Wang et al. FlowPIE: Test-time scientific idea evolution with flow-guided literature exploration, 2026

  55. [55]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. InAdvances in Neural Information Processing Systems, 2022. NeurIPS 2022; arXiv:2201.11903; checked May 10, 2026

  56. [56]

    Novbench: Evaluating large language models on academic paper novelty assessment, 2026

    Wenqing Wu, Yi Zhao, Yuzhuo Wang, Siyou Li, Juexi Shao, Yunfei Long, and Chengzhi Zhang. Novbench: Evaluating large language models on academic paper novelty assessment, 2026. arXiv:2604.11543; checked May 10, 2026

  57. [57]

    Improving scientific hypothesis generation with knowledge grounded large language models, 2024

    Guangzhi Xiong, Eric Xie, Amir Hassan Shariatmadari, Sikun Guo, Stefan Bekiranov, and Aidong Zhang. Improving scientific hypothesis generation with knowledge grounded large language models, 2024. arXiv:2411.02382; checked May 10, 2026

  58. [58]

    The AI scientist-v2: Workshop-level automated scientific discovery via agentic tree search, 2025

    Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The AI scientist-v2: Workshop-level automated scientific discovery via agentic tree search, 2025. arXiv:2504.08066; checked May 10, 2026

  59. [59]

    ReAct: Synergizing reasoning and acting in language models

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. InInternational Conference on Learning Represen- tations, 2023. ICLR 2023; arXiv:2210.03629; checked May 10, 2026. 52

  60. [60]

    Opennovelty: An LLM-powered agentic system for verifiable scholarly novelty assessment, 2026

    Ming Zhang, Kexin Tan, Yueyuan Huang, Yujiong Shen, Chunchun Ma, Li Ju, Xinran Zhang, Yuhui Wang, Wenqing Jing, Jingyi Deng, Huayu Sha, Binze Hu, Jingqi Tong, Changhao Jiang, Yage Geng, Yuankai Ying, Yue Zhang, Zhangyue Yin, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, and Xuanjing Huang. Opennovelty: An LLM-powered agentic system for verifiable scholarly n...

  61. [61]

    Deep ideation: Designing LLM agents to generate novel research ideas on scientific concept network, 2025

    Keyu Zhao, Weiquan Lin, Qirui Zheng, Fengli Xu, and Yong Li. Deep ideation: Designing LLM agents to generate novel research ideas on scientific concept network, 2025. arXiv:2511.02238; checked May 10, 2026

  62. [62]

    From automation to autonomy: A survey on large language models in scientific discovery, 2025

    Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song. From automation to autonomy: A survey on large language models in scientific discovery, 2025. arXiv:2505.13259; checked May 10, 2026

  63. [63]

    Nguyen, Baixuan Xu, Zhaowei Wang, Jiayang Cheng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Tianqing Fang, Yangqiu Song, Ginny Y

    Tianshi Zheng, Kelvin Kiu-Wai Tam, Newt Hue-Nam K. Nguyen, Baixuan Xu, Zhaowei Wang, Jiayang Cheng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, and Simon See. Newtonbench: Benchmarking generalizable scientific law discovery in LLM agents, 2025. arXiv:2510.07172; checked May 10, 2026. 53