REVIEW 28 cited by
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The pace of scientific research, vital for improving human life, is complex, slow, and needs specialized expertise. Meanwhile, novel, impactful research often stems from both a deep understanding of prior work, and a cross-pollination of ideas across domains and fields. To enhance the productivity of researchers, we propose ResearchAgent, which leverages the encyclopedic knowledge and linguistic reasoning capabilities of Large Language Models (LLMs) to assist them in their work. This system automatically defines novel problems, proposes methods and designs experiments, while iteratively refining them based on the feedback from collaborative LLM-powered reviewing agents. Specifically, starting with a core scientific paper, ResearchAgent is augmented not only with relevant publications by connecting information over an academic graph but also entities retrieved from a knowledge store derived from shared underlying concepts mined across numerous papers. Then, mimicking a scientific approach to improving ideas with peer discussions, we leverage multiple LLM-based ReviewingAgents that provide reviews and feedback via iterative revision processes. These reviewing agents are instantiated with human preference-aligned LLMs whose criteria for evaluation are elicited from actual human judgments via LLM prompting. We experimentally validate our ResearchAgent on scientific publications across multiple disciplines, showing its effectiveness in generating novel, clear, and valid ideas based on both human and model-based evaluation results. Our initial foray into AI-mediated scientific research has important implications for the development of future systems aimed at supporting researchers in their ideation and operationalization of novel work.
Forward citations
Cited by 28 Pith papers
-
FARS: A Fully Automated Research System Deployed at Scale
FARS deployed at scale produced 166 AI/ML papers across 67 topics that received 282 structured human reviews indicating some review-worthy outputs alongside recurring failure modes.
-
ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models
Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.
-
Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
Argus demonstrates that a fixed-weight, self-evolving multi-role agentic runtime with verification-gated persistence can achieve competitive benchmark results and retain reusable state across long-horizon tasks.
-
IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
IdeaTrail reverse-synthesizes 1,170 grounded multi-turn agent trajectories from real papers via a Generator–Advisor loop for scientific ideation process supervision.
-
Beyond the Golden Record: Toward a Design Theory for Trustworthy Master Data Management with Self-Sovereign Identity
A design theory is derived for trustworthy master data management based on self-sovereign identity to support reliable, sovereign, and accountable data sharing in data ecosystems.
-
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
In FIRE-Bench's rediscovery test — question-only prompt, methods withheld — the best agent averaged 46.7 F1 and no agent reached 50, with failures concentrated in research planning and conclusion formation.
-
Reinforcement Learning for Machine Learning Engineering Agents
RL-trained Qwen2.5-3B outperforms prompted Claude-3.5-Sonnet and GPT-4o on 12 MLEBench tasks by an average of 22% and 24%, using two targeted RL modifications.
-
HypoChainer: A Collaborative System Combining LLMs and Knowledge Graphs for Hypothesis-Driven Scientific Discovery
In a small user study and two case studies, a hypothesis-chain workflow grounded in knowledge graphs helped biomedical researchers construct and validate hypotheses from machine-learning predictions more effectively t...
-
THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning?
THE-Tree constructs causally-linked semantic evolution trees from surveys and literature, and the authors report improved graph completion, future prediction, and LLM-based paper evaluation.
-
Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation
Doc2Agent automatically converts unstructured REST API documentation into validated, Python-based tools for AI agents, reporting a 55% relative WebArena improvement over direct API calling.
-
EXP-Bench: Can AI Conduct AI Research Experiments?
EXP-Bench is a new benchmark of 461 end-to-end AI research experiments, and leading AI agents complete fewer than 1 percent of them successfully.
-
Harnessing Large Language Models for Scientific Novelty Detection
The authors propose an LLM-distilled retriever trained on rephrased, partial, and incremental idea variants, and show it improves retrieval and novelty detection on two new closed-domain datasets in marketing and NLP.
-
Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
A new benchmark (TruthHypo) and a knowledge-grounded hallucination detector (KnowHD) show that grounding scores can partially select truthful LLM-generated biomedical hypotheses, but the result is at risk from knowled...
-
MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem
MM-Agent, a multi-stage LLM pipeline with a hierarchical modeling method library, is claimed to outperform prior agents and award-winning human solutions on a new 111-problem MCM/ICM-based mathematical modeling benchmark.
-
Automatic Evaluation Metrics for Artificially Generated Scientific Research
A simple title-and-abstract model predicts citation counts better than review scores and outperforms LLM reviewers in matching human review scores, but remains below human consistency.
-
Automated Hypothesis Validation with Agentic Sequential Falsifications
An LLM-agent framework validates free-form hypotheses through sequential falsification experiments aggregated with e-values to control Type-I error.
-
VASP Agent: An Agentic Framework for Autonomous First-principles Calculations
An LLM-driven agent with predefined VASP workflows and parameter-checking tools completes DFT simulation tasks more reliably and accurately than standalone LLMs, with a new 80-task benchmark.
-
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
ResearchPulse extracts motivation-method chains and experimental trends from related papers, rendering them as mind maps and line charts, and releases a 100-cluster benchmark; the reported '7B beats GPT-4o' result is ...
-
Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
WikiHowAgent generates 114,296 simulated teacher-learner conversations from 14,287 WikiHow tutorials and evaluates their pedagogic quality with LLM and human judges.
-
SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion
A multi-agent LLM system called SimuPanel simulates expert panel discussions with personas grounded in public academic sources, and a small evaluation suggests the full reasoning pipeline produces higher LLM-judged di...
-
Simulating Human Behavior with the Psychological-mechanism Agent: Integrating Feeling, Thought, and Action
PSYA combines ALMA emotion layers and the Triple Network Model to make LLM agents behave more human-like and reproduce several classic psychology experiment results.
-
AI-Researcher: Autonomous Scientific Innovation
AI-Researcher runs an end-to-end ML research pipeline with LLM agents, and Scientist-Bench measures how close the resulting papers come to human-authored publications.
-
A Multi-Layered Framework for Modeling Human Biology: From Basic AI Agents to a Full-Body AI Agent
The paper proposes, but does not implement or validate, a multi-agent AI framework for cross-scale modeling of human biology from molecules to whole body, with sketches of metastasis scoring and drug development.
-
AI4Research: A Survey of Artificial Intelligence for Scientific Research
A survey that organizes AI-for-research work into five tasks, comprehension, survey, discovery, writing, and peer review, and compiles associated tools and benchmarks.
-
Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI
The paper argues that integrating cognitive AI and embodied robots into closed-loop Intelligent Science Laboratories is essential for the next leap in automated scientific discovery.
-
A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.
-
Beyond Explainability: The Case for AI Validation
AI governance should shift from explainability to validation as its central regulatory pillar, with a typology of valid-versus-explainable systems.
-
Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research
A perspective article reviews the state of using foundation models for laboratory automation and proposes a roadmap for fully autonomous experiments.
Discussion (0). Sign in to comment.