Pith. sign in

REVIEW 11 cited by

LLM4SR: A Survey on Large Language Models for Scientific Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.04306 v1 pith:UHUYKZA5 submitted 2025-01-08 cs.CL cs.DL

classification cs.CLcs.DL
keywords researchllmsscientificsurveyacrosslanguagelargellm4sr
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, the rapid advancement of Large Language Models (LLMs) has transformed the landscape of scientific research, offering unprecedented support across various stages of the research cycle. This paper presents the first systematic survey dedicated to exploring how LLMs are revolutionizing the scientific research process. We analyze the unique roles LLMs play across four critical stages of research: hypothesis discovery, experiment planning and implementation, scientific writing, and peer reviewing. Our review comprehensively showcases the task-specific methodologies and evaluation benchmarks. By identifying current challenges and proposing future research directions, this survey not only highlights the transformative potential of LLMs, but also aims to inspire and guide researchers and practitioners in leveraging LLMs to advance scientific inquiry. Resources are available at the following repository: https://github.com/du-nlp-lab/LLM4SR

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeepInflation: an AI agent for research and model discovery of inflation

    astro-ph.CO 2026-01 conditional novelty 6.0 of 10

    An LLM agent with symbolic regression finds simple inflation potentials that match target CMB observables, but the outputs are fitted to the targets rather than independently predicted.

  2. EvoVLMA: Evolutionary Vision-Language Model Adaptation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    An LLM-based evolutionary algorithm automatically designs training-free VLM adaptation code, improving few-shot classification accuracy over manually-designed baselines by up to 1.91 points.

  3. PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new multi-turn reinforcement learning framework trains a single LLM to both solve math problems and verify its own solutions, revising only when its verifier finds a mistake.

  4. ScienceMeter: Tracking Scientific Knowledge Updates in Language Models

    cs.CL 2025-05 reject novelty 6.0 of 10

    ScienceMeter evaluates language model knowledge updates across three axes, preservation of old scientific claims, acquisition of new claims, and projection to future findings, and finds all current methods fall short.

  5. Towards Fully Automated Molecular Simulations: Multi-Agent Framework for Simulation Setup and Force Field Extraction

    cs.AI 2025-09 conditional novelty 5.0 of 10

    A multi-agent LLM system generates RASPA simulation inputs and extracts literature force field parameters with moderate to high accuracy on a small set of zeolite tasks.

  6. Data Shift of Object Detection in Autonomous Driving

    cs.RO 2025-08 reject novelty 5.0 of 10

    The abstract's claim of superior BDD100K object-detection performance has no supporting content in the full text, which is a different paper whose LLM-designed CMOEA modules beat 11 baselines on benchmarks the modules...

  7. A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection

    cs.CR 2025-08 conditional novelty 5.0 of 10

    RTST, a two-agent moderator with an explainable Behavior ledger and per-prompt weight updates, reduced attack success rate from 12-63% to 0-17% on three jailbreak benchmarks with Gemini 2.5 Flash.

  8. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  9. Conversational AI for Rapid Scientific Prototyping: A Case Study on ESA's ELOPE Competition

    cs.AI 2026-01 conditional novelty 4.0 of 10

    One engineer paired with ChatGPT and reached second place in ESA's ELOPE competition in about one week of work; the paper draws best-practice lessons from that experience.

  10. Accelerating Scientific Discovery with Multi-Document Summarization of Impact-Ranked Papers

    cs.DL 2025-08 conditional novelty 4.0 of 10

    The authors add an LLM-powered summarization tool to the BIP! Finder search engine that generates cited, concise or review-style summaries of impact-ranked search results.

  11. How Far Are AI Scientists from Changing the World?

    cs.AI 2025-07 conditional novelty 4.0 of 10

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

Pith tools