Pith. sign in

REVIEW 3 major objections 7 minor 35 cited by

AI4Research: A Survey of Artificial Intelligence for Scientific Research

T0 review · 3 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A new survey maps AI for research onto five stage-level tasks and composes them into one pipeline.

desk verdict Useful organizational survey; the formal composition is decorative and the taxonomy is the real contribution. read the letter →

arxiv 2507.01903 v2 pith:6ESIYSBX submitted 2025-07-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords AI4Researchlargelanguagemodelsscientificdiscoveryacademicsurveywritingpeerreviewcomprehensionresearchtaxonomy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to establish that AI for Research (AI4Research) is a field with a single spine: five task families — scientific comprehension, academic survey, scientific discovery, academic writing, and peer reviewing — that can be composed into one pipeline running from a research query to a reviewed publication. The authors argue this unified frame is what has been missing: earlier surveys focused on discovery and writing, so researchers lacked a map of the full lifecycle and the resources at each stage. If the frame holds, a newcomer can locate any AI research tool or benchmark in one of five slots, and progress at any stage can be measured against the pipeline as a whole. The paper couples the taxonomy with a formal composition (Eq. 1-2) and a per-stage resource compilation spanning the natural, applied, and social sciences.

What carries the argument

The load-bearing object is the five-part taxonomy itself, made formal by the composition equation $A = A_{PR} \circ A_{AW} \circ A_{SD} \circ A_{AS} \circ A_{SC}$ (Eq. 1), where $\circ$ is function composition and each $A_i$ is the AI model tailored to one research task: comprehension ($A_{SC}$), survey ($A_{AS}$), discovery ($A_{SD}$), writing ($A_{AW}$), and peer review ($A_{PR}$). Applied to a research query $q$ (Eq. 2), the composed system produces a publication $A(q)$, and Eq. 3 states the objective as maximizing efficiency, performance, and innovation. The taxonomy does the work of the paper: it is the organizing frame that lets every surveyed system be placed in the lifecycle, gives each stage a definition in terms of input-output functions (Eqs. 4-13), and turns 'AI4Research' from a slogan into a pipeline with replaceable stages.

What would settle it

Run an independent, criteria-driven sweep of recent AI-for-research systems (for example, fixed search strings across a defined set of venues and years) and ask two questions: does every system fall into exactly one of the five taxonomy branches, and do the survey's comparison tables (Tables 2-5, sourced from other papers) reproduce under re-annotation? The central claim weakens if a substantial share of systems straddle branches or fit none, or if the reproduced numbers do not match the cited sources.

Watch

Extended reading notes

Core claim

The paper's central claim is that the field of AI for research can be organized — and should be understood — as five task families that together form a single pipeline: AI for Scientific Comprehension (extracting knowledge from a paper's text, tables, and charts), AI for Academic Survey (retrieving and synthesizing many papers into overviews and related-work sections), AI for Scientific Discovery (idea mining, novelty assessment, theory analysis, and experiment conduction), AI for Academic Writing (assisting or fully automating manuscript production), and AI for Academic Peer Reviewing (pre-review, in-review, and post-review stages). It formalizes this as the functional composition $A = A_{PR} \circ A_{AW} \circ A_{SD} \circ A_{AS} \circ A_{SC}$, so that a research query $q$ flows through comprehension, survey, discovery, writing, and review to yield a reviewed publication, and it states the system goal as maximizing efficiency, performance, and innovation (Eq. 3). On this basis the paper positions itself as the missing survey that spans the whole lifecycle: earlier surveys, it argues, covered only discovery and writing under the banner of AI4Science, whereas AI4Research includes scientific comprehension, academic survey, and peer review as first-class stages. It then identifies priority gaps — the rigor and scalability of automated experiments and societal impact — and compiles per-stage tools, datasets, and benchmarks across natural, applied, and social sciences.

Load-bearing premise

The whole frame rests on the assumption that the five-way taxonomy and its descriptions of individual systems are faithful to the underlying literature; because the paper sets no systematic search or inclusion criteria, a biased or inaccurate selection would leave the unified perspective unsupported even if every individual summary were right.

Editorial extensions

If this is right

  • A researcher entering AI4Research can locate any existing system — from PaperQA2 to the AI Scientist to DeepReview — in one of five task families and see which stage of the lifecycle it automates.
  • Stage-level progress becomes measurable: because the framework composes the five stages, improving one module (for example, retrieval-backed surveys) can be evaluated by its effect on downstream discovery and writing quality.
  • The paper's identified gaps — rigor and scalability of automated experiments, plus societal impact — become concrete research targets rather than vague cautions.
  • The resource lists (benchmarks, datasets, tools per stage) lower the cost of entry for building or evaluating AI4Research systems.
  • The equations give a formal target: an AI4Research system is good insofar as it maximizes the efficiency, performance, and innovation of the publication it produces.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next step the paper does not take is to use its composition equation as a benchmark design, fixing every stage except the one under test, so that gains in, say, idea mining can be isolated from losses in writing or review.
  • The five-stage pipeline reads as an engineering blueprint: if stages are connected by standardized interfaces (a query, a survey, an idea, a manuscript, a review), teams could upgrade modules independently instead of rebuilding whole 'AI scientist' systems.
  • Because the survey sets no systematic inclusion criteria, a reader should treat its coverage as a curated map rather than a census; the same taxonomy could later be tested by a meta-analysis with explicit search strings and inter-annotator agreement.
  • The AI4Science versus AI4Research split implies a division of labour: domain-specific discovery systems plug into a broader research-workflow shell, which suggests integration (API-style composition) rather than competition between the two lines of work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper presents a survey of AI for research (AI4Research), organizing the field into five task areas: AI for Scientific Comprehension, AI for Academic Survey, AI for Scientific Discovery, AI for Academic Writing, and AI for Academic Peer Reviewing. It introduces a formal composition of these task modules in Eqs. (1)-(2), defines per-module objectives in Eqs. (4)-(13), contrasts AI4Research with AI4Science, surveys methods and applications across natural, applied, and social sciences, compiles tools and datasets in Section 9, and proposes future directions in Section 10. The paper also reproduces several comparison tables from prior benchmark papers.

Significance. If the taxonomy and resource map are accurate, this survey is useful as an organizing frame and entry point for a rapidly growing literature. Its strengths include the breadth of the reference list, a generally sensible decomposition of the research lifecycle into five tasks, transparent attribution of benchmark tables to the original papers, and a substantial collection of tools, datasets, and applications. The formal composition in Eqs. (1)-(2), however, is not mathematically well-defined, and the optimization objectives in Eqs. (3)-(13) are presented without definitions or a probabilistic model; these formal elements should not be presented as a unified mathematical foundation without substantial repair or reframing.

major comments (3)
  1. [§2, Eqs. (1)-(2) and (4)-(13)] The composition A = APR ∘ AAW ∘ ASD ∘ AAS ∘ ASC is not type-correct under the definitions given in Eqs. (4)-(13). ASC in Eq. (4) is described as taking documents DSC and producing knowledge K; AAS in Eq. (6) takes survey requirements RAS and produces a survey S; ASD in Eq. (8) produces innovations I; AAW in Eq. (10) produces a manuscript M; and APR in Eq. (12) produces a review R. The codomain of each stage is not the domain of the next stage, and the query q in Eq. (2) is not obviously in the domain of ASC, which expects documents. The equations therefore describe a schematic pipeline, not a function composition. Please either provide explicit domain/codomain types for each module and show that they chain, or remove the claim that Eqs. (1)-(2) provide a formal unified perspective.
  2. [§2, Eqs. (3), (5), (7), (9), (11), (13)] The stated objectives are not well-defined. In Eq. (3), η(·), α(·), and τ(·) are not defined. In Eqs. (5)-(13), the quantities Coherence, Coverage, Relevance, Clarity, Novelty, Validity, Significance, Consistency, Readability, Compliance, Correctness, Helpfulness, and the expectations over K ∼ ASC, S ∼ AAS, I ∼ ASD, M ∼ AAW, and R ∼ APR have no specified definitions or probability models. As written, these expressions cannot be evaluated or optimized, so they do not provide a formal foundation for the proposed unified perspective. Either define these terms operationally or present them as informal schematics rather than as formal objectives.
  3. [Abstract and §1 (Introduction)] The paper claims to fill the absence of a comprehensive survey on AI4Research, but it does not report any literature search methodology. There is no statement of databases queried, time span covered, inclusion or exclusion criteria, keyword strategy, or screening process. Without such information, the comprehensiveness claim is not verifiable and the selection of cited works may be biased. Please add a methodology paragraph or, if the survey is intentionally selective, temper the comprehensiveness claim accordingly.
minor comments (7)
  1. [§2, first paragraph] The text says 'we identify six core capabilities' but then lists only five: Scientific Comprehension, Academic Survey, Scientific Discovery, Academic Writing, and Academic Peer Reviewing. Change 'six' to 'five' or add the missing item.
  2. [§2.1.1, Eq. (4)] The symbol K appears both as the input and the output of ASC in Eq. (4), and the actual input documents DSC are not represented in the left-hand side. Clarify the input/output notation; the same ambiguity affects Eqs. (6), (8), (10), and (12).
  3. [§2, after Eq. (2)] There is a typo: 'Furhter' should be 'Further'.
  4. [§3.1.1] 'give a manually created question' should be 'given a manually created question'.
  5. [Figure 2] In the Semi-Automatic Academic Writing list, TikZero [55] appears twice; remove the duplicate entry.
  6. [Figure 3] The axis label 'Level of Automation ext' appears garbled, and the percentage '70%' is unexplained. Please fix the figure text.
  7. [§5.5] The statement that papers generated through Zochi 'have even been accepted by ACL 2025' is a strong, specific claim without a citation or venue details. Please add evidence or soften the claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey’s taxonomy and formal definitions are organizational stipulations, not derived predictions or self-citation-forced results.

full rationale

This is a survey paper whose central claim is that the field of AI4Research can be organized into five tasks: Scientific Comprehension, Academic Survey, Scientific Discovery, Academic Writing, and Academic Peer Reviewing. That claim is presented as an explicit taxonomy (“We first introduce a systematic taxonomy to classify five mainstream tasks in AI4Research”), not as a theorem derived from independent premises. The formal equations in Section 2 are definitions of module compositions and objective functions: Eq. 1 states a functional composition, Eq. 2 applies it to a query, and Eqs. 4–13 define each module’s input/output schematically. None of these equations is fitted to data, none of them predicts an observable quantity, and no empirical result is presented as following from the equations. The comparison tables in Sections 4, 5, and 7 are explicitly attributed to external sources (Yan et al., Ruan et al., Chen et al., Shin et al.), so they are not internally generated inputs renamed as outputs. The paper’s emphasis on the absence of a comprehensive survey is supported by a review of existing surveys in Section 11; while some cited works are by the present authors, the five-way organizational frame is not justified by any self-citation chain or uniqueness theorem. A separate formal concern is that the signatures in Eq. 1 do not strictly type-check (e.g., ASC outputs knowledge while AAS expects survey requirements), but this is a consistency or presentation issue, not circularity: the composition is a schematic definition rather than a derived result. The paper is self-contained as an organizational and resource survey, so the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or new entities are introduced. The central claim depends instead on the accuracy of external literature and on the adopted taxonomy.

assumptions (2)
  • ad hoc to paper The five-task taxonomy partitions the research lifecycle.
    This is the paper's organizing contribution; it is asserted rather than derived. It appears in Section 1, Section 2, and Figure 1.
  • domain assumption The surveyed systems are described accurately by the cited sources.
    The survey's synthesis depends on the reliability of its references; no independent verification is provided for most entries.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI4Research: A Survey of Artificial Intelligence for Scientific Research." pith.science (2026). https://pith.science/paper/6ESIYSBX

@misc{pith2026250701903,
  author       = {Pith},
  title        = {Pith review of: AI4Research: A Survey of Artificial Intelligence for Scientific Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6ESIYSBX}},
  note         = {Machine review of arXiv:2507.01903}
}
read the original abstract

Recent advancements in artificial intelligence (AI), particularly in large language models (LLMs) such as OpenAI-o1 and DeepSeek-R1, have demonstrated remarkable capabilities in complex domains such as logical reasoning and experimental coding. Motivated by these advancements, numerous studies have explored the application of AI in the innovation process, particularly in the context of scientific research. These AI technologies primarily aim to develop systems that can autonomously conduct research processes across a wide range of scientific disciplines. Despite these significant strides, a comprehensive survey on AI for Research (AI4Research) remains absent, which hampers our understanding and impedes further development in this field. To address this gap, we present a comprehensive survey and offer a unified perspective on AI4Research. Specifically, the main contributions of our work are as follows: (1) Systematic taxonomy: We first introduce a systematic taxonomy to classify five mainstream tasks in AI4Research. (2) New frontiers: Then, we identify key research gaps and highlight promising future directions, focusing on the rigor and scalability of automated experiments, as well as the societal impact. (3) Abundant applications and resources: Finally, we compile a wealth of resources, including relevant multidisciplinary applications, data corpora, and tools. We hope our work will provide the research community with quick access to these resources and stimulate innovative breakthroughs in AI4Research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 35 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

    cs.AI 2026-05 unverdicted novelty 8.0 of 10

    SciIntegrity-Bench shows state-of-the-art LLMs violate academic integrity in 34.2% of dilemmatic scenarios, primarily by fabricating data rather than refusing impossible tasks.

  2. SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

    cs.AI 2026-05 unverdicted novelty 8.0 of 10

    SciIntegrity-Bench shows seven LLMs exhibit a 34.2% integrity failure rate in dilemmatic scenarios, with all models fabricating synthetic data in missing-data cases and an intrinsic completion bias persisting after pr...

  3. Hidden-State Privacy Has an Empty Middle

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    No Gaussian release fills the moderate utility-moderate privacy middle for hidden states; a Fisher lower bound rules out uniform safety in the class, with only edge diagonal mechanisms succeeding and new split-memory ...

  4. Hidden-State Privacy Has an Empty Middle

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    No full-rank Gaussian release of single-layer hidden states achieves moderate utility and moderate privacy against adaptive Mahalanobis retrieval; only edge mechanisms or architectural redesign work.

  5. Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    DataPRM is a new process reward model for data analysis agents that detects silent errors via environment interaction and ternary rewards, yielding 7-11% gains on benchmarks and further RL improvements.

  6. Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    DataPRM is an environment-aware generative process reward model that improves LLM data analysis agents by 7-11% on benchmarks via active verification and reflection-aware ternary rewards.

  7. Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends

    cs.CR 2025-08 accept novelty 7.0 of 10

    A survey of LLM copyright protection that unifies text watermarking, model watermarking, and model fingerprinting while presenting new coverage of fingerprint transfer and removal.

  8. AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A benchmark of 1,500 expert-annotated ablation study designs from 807 NLP papers shows frontier LLMs underperform human experts and that LLM-as-a-judge evaluations correlate weakly with human judgments.

  9. RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

    cs.SE 2026-07 conditional novelty 6.0 of 10

    A controlled benchmark shows LLM agents can sometimes discover better training-data strategies through feedback, but their improvements are fragile and usually not sustained.

  10. Judgment-Grounded Expansion for Peer Review Generation

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Formalizes judgment-grounded expansion as a human-AI collaborative task for peer review generation, supported by a user study and conformal prediction methods for scalable evaluation.

  11. Bayesian Adaptation Gym: A Benchmark for the Bayesian Low-Rank Adaptation of Multi-Modal Language Models

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Introduces BAG, an open-source benchmark suite with baselines, datasets, and tasks for assessing Bayesian low-rank adaptation of multi-modal language models on calibration, robustness, and decision-making under uncertainty.

  12. LLM-AutoSciLab: Closed-Loop Scientific Discovery via Active Experimentation with LLMs

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    LLM-AutoSciLab proposes an LLM-driven closed-loop system for hypothesis generation and adaptive experiment selection that reports higher accuracy and 2-5x better sample efficiency than baselines on new chemistry and g...

  13. The Scaling Laws of Skills in LLM Agent Systems

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Empirical analysis across 15 LLMs and 1,141 skills identifies a logarithmic routing decay law and a multiplicative execution law coupled by a single fitted slope parameter b that enables targeted library optimizations...

  14. SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    SciHorizon-DataEVA is a hierarchical multi-agent system that applies Sci-TQA2 principles to assess AI-readiness of heterogeneous scientific data through dynamic evaluation specifications and adaptive tool use.

  15. When AI reviews science: Can we trust the referee?

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    AI peer review systems are vulnerable to prompt injections, prestige biases, assertion strength effects, and contextual poisoning, as demonstrated by a new attack taxonomy and causal experiments on real conference sub...

  16. How Researchers Navigate Accountability, Transparency, and Trust When Using AI Tools in Early-Stage Research: A Think-Aloud Study

    cs.CY 2026-04 unverdicted novelty 6.0 of 10

    A think-aloud study reveals that AI tools in early research misrepresent uncertainty, obscure provenance, and create fragile trust, leading researchers to develop compensatory strategies to preserve scholarly judgment.

  17. OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    OMIBench benchmark reveals that current LVLMs achieve at most 50% on Olympiad problems requiring reasoning across multiple images.

  18. MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    MeasHalu reduces LLM hallucinations in scientific measurement extraction via a fine-grained taxonomy, reasoning-aware fine-tuning, and progressive rewards, improving accuracy on the MeasEval benchmark.

  19. AI Can Learn Scientific Taste

    cs.CL 2026-03 conditional novelty 6.0 of 10

    Reinforcement learning on citation-preference pairs teaches a model to predict which papers will be cited more and to propose ideas that LLM judges rate as likely to be cited more—but "taste" here means citation impact.

  20. Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems

    cs.NE 2026-07 accept novelty 5.0 of 10

    Evolutionary intelligence reframes evolutionary computation as cumulative scientific discovery by retaining search trajectories, failures, and lineages across cycles.

  21. Can AI Review Improve Paper Drafting? An Empirical Study on 20 Computer Architecture Submissions

    cs.AI 2026-05 unverdicted novelty 5.0 of 10

    An empirical study on 20 architecture papers finds AI reviews capture a significant fraction of human-raised issues while also surfacing additional ones, using a released tool that clusters AI comments for comparison.

  22. Heterogeneous Scientific Foundation Model Collaboration

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Eywa enables language-based agentic AI systems to collaborate with specialized scientific foundation models for improved performance on structured data tasks.

  23. AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    AblateCell reproduces baselines in three single-cell perturbation repositories with 88.9% success and recovers ground-truth critical components with 93.3% accuracy via closed-loop ablation.

  24. Emergent Social Intelligence Risks in Generative Multi-Agent Systems

    cs.MA 2026-03 unverdicted novelty 5.0 of 10

    Generative multi-agent systems exhibit emergent collusion and conformity behaviors that cannot be prevented by existing agent-level safeguards.

  25. Deep Researcher with Test-Time Diffusion

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A test-time 'denoising' loop that repeatedly revises a draft report using fresh web retrieval, plus a component-wise self-evolution step, beats existing deep research agents on several benchmarks.

  26. ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Interleaving key video frames into step-by-step reasoning improves video question answering by 1.7 to 5.5 points over text-only chain-of-thought on a new self-built benchmark.

  27. Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

    cs.AI 2025-03 unverdicted novelty 5.0 of 10

    The paper unifies perspectives on Long CoT in reasoning LLMs by introducing a taxonomy, detailing characteristics of deep reasoning and reflection, and discussing emergence phenomena and future directions.

  28. Hephaestus: Toward a Cybersecurity AI Scientist

    cs.CR 2026-06 unverdicted novelty 4.0 of 10

    The paper proposes the Cybersecurity AI Scientist as a modular multi-agent architecture for automating cybersecurity research, distinguished by its focus on non-stationary threats and anchored in a four-zeros risk-tru...

  29. AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    A survey organizing AI-powered research automation into five workflow stages, defining AutoResearch and Vibe Research, and proposing five evaluation dimensions while noting domain-conditioned limits on autonomy.

  30. SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    SciAtlas builds a large-scale multi-disciplinary academic knowledge graph and a neuro-symbolic retrieval system to support automated scientific research tasks such as literature review and idea positioning.

  31. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 conditional novelty 4.0 of 10

    AI can generate research artifacts faster than it can verify them, so across all eight lifecycle stages the credible deployment mode is human-governed collaboration rather than full autonomy.

  32. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  33. SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

    cs.AI 2026-04 unverdicted novelty 4.0 of 10

    SciHorizon-DataEVA is a multi-agent system that applies Sci-TQA2 principles across four dimensions to assess AI-readiness of heterogeneous scientific data via dynamic profiling and self-correcting evaluation workflows.

  34. A Survey of Context Engineering for Large Language Models

    cs.CL 2025-07 accept novelty 4.0 of 10

    The survey organizes Context Engineering into retrieval, processing, management, and integrated systems like RAG and multi-agent setups while identifying an asymmetry where LLMs handle complex inputs well but struggle...

  35. Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator

    cs.DL 2025-07 unverdicted novelty 4.0 of 10

    The paper proposes a four-role framework for LLMs in scientific innovation and reviews methods, benchmarks, and limitations across Assistant, Collaborator, Scientist, and Evaluator roles.

Reference graph

Works this paper leans on

300 extracted references · 26 canonical work pages · cited by 30 Pith papers

  1. [1]

    Phi-4 technical report.arXiv preprint arXiv:2412.08905, 2024

    Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. Phi-4 technical report.arXiv preprint arXiv:2412.08905, 2024

  2. [2]

    Accurate structure prediction of biomolecular interactions with alphafold 3.Nature, 630(8016):493–500, May 2024

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ron- neberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. Accurate structure prediction of biomolecular interactions with alphafold 3.Nature, 630(8016):493–500, May 2024

  3. [3]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  4. [4]

    Ask, retrieve, summarize: A modular pipeline for scientific literature summarization.arXiv preprint arXiv:2505.16349, 2025

    Pierre Achkar, Tim Gollub, and Martin Potthast. Ask, retrieve, summarize: A modular pipeline for scientific literature summarization.arXiv preprint arXiv:2505.16349, 2025

  5. [5]

    Efficient bayesian learningcurveextrapolationusingprior-datafittednetworks

    Steven Adriaensen, Herilalaina Rakotoarison, Samuel Müller, and Frank Hutter. Efficient bayesian learningcurveextrapolationusingprior-datafittednetworks. AdvancesinNeuralInformationProcessing Systems, 36:19858–19886, Dec 2023

  6. [6]

    Llm4grn: Discovering causal gene regulatory networks with llms–evaluation through synthetic data generation.arXiv preprint arXiv:2410.15828, 2024

    Tejumade Afonja, Ivaxi Sheth, Ruta Binkyte, Waqar Hanif, Thomas Ulas, Matthias Becker, and Mario Fritz. Llm4grn: Discovering causal gene regulatory networks with llms–evaluation through synthetic data generation.arXiv preprint arXiv:2410.15828, 2024

  7. [7]

    Litllm: A toolkit for scientific literature review.arXiv preprint arXiv:2402.01788, 2024

    Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H Laradji, Krishnamurthy DJ Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. Litllm: A toolkit for scientific literature review.arXiv preprint arXiv:2402.01788, 2024

  8. [8]

    Llms for literature review: Are we there yet?arXiv preprint arXiv:2412.15249, 2024

    Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H Laradji, Krishnamurthy DJ Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. Llms for literature review: Are we there yet?arXiv preprint arXiv:2412.15249, 2024

Show all 300 references
  1. [9]

    Laradji, Krishnamurthy Dj Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal

    Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H. Laradji, Krishnamurthy Dj Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. LitLLMs, LLMs for literature review: Are we there yet? Transactions on Machine Learning Research, Apr 2025. ISSN 2835-8856. URL https://...

  2. [10]

    Artificial intelligence and scientific discovery: A model of prioritized search.Research Policy, 53(5):104989, Jun 2024

    Ajay Agrawal, John McHale, and Alexander Oettl. Artificial intelligence and scientific discovery: A model of prioritized search.Research Policy, 53(5):104989, Jun 2024

  3. [11]

    Ml-gap: machine learning-enhanced genomic analysis pipeline using autoencoders and data augmentation.Frontiers in Genetics, 15:1442759, Sep 2024

    Melih Agraz, Dincer Goksuluk, Peng Zhang, Bum-Rak Choi, Richard T Clements, Gaurav Choudhary, and George Em Karniadakis. Ml-gap: machine learning-enhanced genomic analysis pipeline using autoencoders and data augmentation.Frontiers in Genetics, 15:1442759, Sep 2024

  4. [12]

    Zochi technical report, Mar 2025

    Intology AI. Zochi technical report, Mar 2025. URLhttps://github.com/IntologyAI/Zochi/ blob/main/Zochi_Technical_Report.pdf. Zochi Technical Report

  5. [13]

    Autonomousmachinelearning-basedpeerreviewer selection system

    NurmukhammedAitymbetovandDimitriosZorbas. Autonomousmachinelearning-basedpeerreviewer selection system. InProceedings of the 31st International Conference on Computational Linguistics: System Demonstrations, pages 199–207, Jan 2025

  6. [14]

    How deep do large language models internalize scientific literature and citation practices? arXiv preprint arXiv:2504.02767, 2025

    Andres Algaba, Vincent Holst, Floriano Tori, Melika Mobini, Brecht Verbeken, Sylvia Wenmackers, and Vincent Ginis. How deep do large language models internalize scientific literature and citation practices? arXiv preprint arXiv:2504.02767, 2025

  7. [15]

    A survey on hypothesis generation for scientific discovery in the era of large language models.arXiv preprint arXiv:2504.05496, 2025

    Atilla Kaan Alkan, Shashwat Sourav, Maja Jablonska, Simone Astarita, Rishabh Chakrabarty, Nikhil Garuda, Pranav Khetarpal, Maciej Pióro, Dimitrios Tanoglidis, Kartheik G Iyer, et al. A survey on hypothesis generation for scientific discovery in the era of large language models...

  8. [16]

    aedfact: Scientific fact-checking made easier via semi-automatic discovery of relevant expert opinions.arXiv preprint arXiv:2305.07796, 2023

    Enes Altuncu, Jason RC Nurse, Meryem Bagriacik, Sophie Kaleba, Haiyue Yuan, Lisa Bonheme, and Shujun Li. aedfact: Scientific fact-checking made easier via semi-automatic discovery of relevant expert opinions.arXiv preprint arXiv:2305.07796, 2023

  9. [17]

    Zero-shot scientific claim verification using llms and citation text

    Carlos Alvarez, Maxwell Bennett, and Lucy Lu Wang. Zero-shot scientific claim verification using llms and citation text. InProceedings of the Fourth Workshop on Scholarly Document Processing (SDP 2024), pages 269–276, Aug 2024

  10. [18]

    Policy advice and best practices on bias and fairness in ai.Ethics and Information Technology, 26(2):31, Apr 2024

    Jose M Alvarez, Alejandra Bringas Colmenarejo, Alaa Elobaid, Simone Fabbrizzi, Miriam Fahimi, Antonio Ferrara, Siamak Ghodsi, Carlos Mougan, Ioanna Papageorgiou, Paula Reyero, et al. Policy advice and best practices on bias and fairness in ai.Ethics and Information Technology,...

  11. [19]

    Languages are still a major barrier to global science.PLoS biology, 14(12):e2000933, Dec 2016

    Tatsuya Amano, Juan P González-Varo, and William J Sutherland. Languages are still a major barrier to global science.PLoS biology, 14(12):e2000933, Dec 2016

  12. [20]

    Ten tips for overcoming language barriers in science.Nature Human Behaviour, 5(9):1119–1122, Jul 2021

    Tatsuya Amano, Clarissa Rios Rojas, Yap Boum II, Margarita Calvo, and Biswapriya B Misra. Ten tips for overcoming language barriers in science.Nature Human Behaviour, 5(9):1119–1122, Jul 2021

  13. [21]

    Homogenization effects of large language models on human creative ideation

    Barrett R Anderson, Jash Hemant Shah, and Max Kreminski. Homogenization effects of large language models on human creative ideation. InProceedings of the 16th conference on creativity & cognition, pages 413–425, Jun 2024

  14. [22]

    Closed-loop transfer enables artificial intelligence to yield chemical knowledge.Nature, 633(8029):351–358, 2024

    Nicholas H Angello, David M Friday, Changhyun Hwang, Seungjoo Yi, Austin H Cheng, Tiara C Torres-Flores, Edward R Jira, Wesley Wang, Alán Aspuru-Guzik, Martin D Burke, et al. Closed-loop transfer enables artificial intelligence to yield chemical knowledge.Nature, 633(8029):351...

  15. [23]

    Transforming science labs into automated factories of discovery.Science Robotics, 9(95):eadm6991, 2024

    Angelos Angelopoulos, James F Cahoon, and Ron Alterovitz. Transforming science labs into automated factories of discovery.Science Robotics, 9(95):eadm6991, 2024

  16. [24]

    The claude 3 model family: Opus, sonnet, haiku

    AI Anthropic. The claude 3 model family: Opus, sonnet, haiku. Claude-3 Model Card, Mar 2024

  17. [25]

    Meta-designing quantum experiments with language models.arXiv preprint arXiv:2406.02470, 2024

    Sören Arlt, Haonan Duan, Felix Li, Sang Michael Xie, Yuhuai Wu, and Mario Krenn. Meta-designing quantum experiments with language models.arXiv preprint arXiv:2406.02470, 2024

  18. [26]

    Openscholar: Synthesizing scientific literature with retrieval-augmented lms.arXiv preprint arXiv:2411.14199, 2024

    Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D’arcy, et al. Openscholar: Synthesizing scientific literature with retrieval-augmented lms.arXiv preprint arXiv:2411.14199, 2024

  19. [27]

    How ai ideas affect the creativity, diversity, and evolution of human ideas: evidence from a large, dynamic experiment

    Joshua Ashkinaze, Julia Mendelsohn, Li Qiwei, Ceren Budak, and Eric Gilbert. How ai ideas affect the creativity, diversity, and evolution of human ideas: evidence from a large, dynamic experiment. arXiv preprint arXiv:2401.13481, 2024

  20. [28]

    The mighty torr: A benchmark for table reasoning and robustness.arXiv preprint arXiv:2502.19412, 2025

    Shir Ashury-Tahan, Yifan Mai, Ariel Gera, Yotam Perlitz, Asaf Yehudai, Elron Bandel, Leshem Choshen, Eyal Shnarch, Percy Liang, Michal Shmueli-Scheuer, et al. The mighty torr: A benchmark for table reasoning and robustness.arXiv preprint arXiv:2502.19412, 2025

  21. [29]

    gpt-researcher, May 2023

    Assafelovic. gpt-researcher, May 2023. URL https://github.com/assafelovic/ gpt-researcher. gpt-researcher

  22. [30]

    Generating fact checking explanations

    Pepa Atanasova. Generating fact checking explanations. InAccountable and Explainable Methods for Complex Reasoning over Text, pages 83–103. Springer, Apr 2024

  23. [31]

    Personalized graph-based retrieval for large language models.arXiv preprint arXiv:2501.02157, 2025

    Steven Au, Cameron J Dimacali, Ojasmitha Pedirappagari, Namyong Park, Franck Dernoncourt, Yu Wang, Nikos Kanakaris, Hanieh Deilamsalehy, Ryan A Rossi, and Nesreen K Ahmed. Personalized graph-based retrieval for large language models.arXiv preprint arXiv:2501.02157, 2025

  24. [32]

    The sciqa scientific question answering benchmark for scholarly knowledge.Scientific Reports, 13(1):7240, May 2023

    Sören Auer, Dante AC Barone, Cassiano Bartz, Eduardo G Cortes, Mohamad Yaser Jaradeh, Oliver Karras, Manolis Koubarakis, Dmitry Mouromtsev, Dmitrii Pliukhin, Daniil Radyush, et al. The sciqa scientific question answering benchmark for scholarly knowledge.Scientific Reports, 13...

  25. [33]

    Self-driving labs are the new ai asset.Axios, Aug 2024

    Axios. Self-driving labs are the new ai asset.Axios, Aug 2024. URLhttps://www.axios.com/ 2024/08/09/ai-self-driving-science-labs-research

  26. [34]

    Robustness evaluation of offline reinforcement learning for robot control against action perturbations.arXiv preprint arXiv:2412.18781, 2024

    Shingo Ayabe, Takuto Otomo, Hiroshi Kera, and Kazuhiko Kawamoto. Robustness evaluation of offline reinforcement learning for robot control against action perturbations.arXiv preprint arXiv:2412.18781, 2024

  27. [35]

    Futuregen: Llm-rag approach to generate the future work of scientific article.arXiv preprint arXiv:2503.16561, 2025

    Ibrahim Al Azher, Miftahul Jannat Mokarrama, Zhishuai Guo, Sagnik Ray Choudhury, and Hamed Alhoori. Futuregen: Llm-rag approach to generate the future work of scientific article.arXiv preprint arXiv:2503.16561, 2025

  28. [36]

    Researchagent: Itera- tive research idea generation over scientific literature with large language models.arXiv preprint arXiv:2404.07738, 2024

    Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. Researchagent: Itera- tive research idea generation over scientific literature with large language models.arXiv preprint arXiv:2404.07738, 2024

  29. [37]

    Scientific paper recommendation: A survey.Ieee Access, 7:9324–9339, Jan 2019

    Xiaomei Bai, Mengyang Wang, Ivan Lee, Zhuo Yang, Xiangjie Kong, and Feng Xia. Scientific paper recommendation: A survey.Ieee Access, 7:9324–9339, Jan 2019. 51

  30. [38]

    Language models surface the unwritten code of science and society.arXiv preprint arXiv:2505.18942, 2025

    Honglin Bao, Siyang Wu, Jiwoong Choi, Yingrong Mao, and James A Evans. Language models surface the unwritten code of science and society.arXiv preprint arXiv:2505.18942, 2025

  31. [39]

    Piors: Personalized intelligent outpatient reception based on large language model with multi-agents medical scenario simulation.arXiv preprint arXiv:2411.13902, 2024

    Zhijie Bao, Qingyun Liu, Ying Guo, Zhengqiang Ye, Jun Shen, Shirong Xie, Jiajie Peng, Xuanjing Huang, and Zhongyu Wei. Piors: Personalized intelligent outpatient reception based on large language model with multi-agents medical scenario simulation.arXiv preprint arXiv:2411.13902, 2024

  32. [40]

    Automated machine learning: past, present and future.Artificial intelligence review, 57(5): 122, Apr 2024

    Mitra Baratchi, Can Wang, Steffen Limmer, Jan N van Rijn, Holger Hoos, Thomas Bäck, and Markus Olhofer. Automated machine learning: past, present and future.Artificial intelligence review, 57(5): 122, Apr 2024

  33. [41]

    Google deepmind’s ai dreamed up 380,000 new materials

    Gregory Barber. Google deepmind’s ai dreamed up 380,000 new materials. the next chal- lenge is making them. WIRED, Nov 2023. URL https://www.wired.com/story/ an-ai-dreamed-up-380000-new-materials-the-next-challenge-is-making-them/

  34. [42]

    Eight years of automl: categorisation, review and trends.Knowledge and Information Systems, 65(12):5097–5149, Aug 2023

    Rafael Barbudo, Sebastián Ventura, and José Raúl Romero. Eight years of automl: categorisation, review and trends.Knowledge and Information Systems, 65(12):5097–5149, Aug 2023

  35. [43]

    Large physics models: Towards a collaborative approach with large language models and foundation models.arXiv preprint arXiv:2501.05382, 2025

    Kristian G Barman, Sascha Caron, Emily Sullivan, Henk W de Regt, Roberto Ruiz de Austri, Mieke Boon, Michael Färber, Stefan Fröse, Faegheh Hasibi, Andreas Ipp, et al. Large physics models: Towards a collaborative approach with large language models and foundation models.arXiv ...

  36. [44]

    Annabel R Basford, Aaron H Bernardino, Paula CP Teeuwen, Benjamin D Egleston, Joshua Humphreys, Kim E Jelfs, Jonathan R Nitschke, Imogen A Riddell, and Rebecca L Greenaway. Development of an automated workflow for screening the assembly and host–guest behavior of metal-organic...

  37. [45]

    The quality assist: A technology-assisted peer review based on citation functions to predict the paper quality.IEEE Access, 10:126815–126831, Dec 2022

    Setio Basuki and Masatoshi Tsuchiya. The quality assist: A technology-assisted peer review based on citation functions to predict the paper quality.IEEE Access, 10:126815–126831, Dec 2022

  38. [46]

    Bruno C Batista, SV Amrutha, Jie Yan, Beni B Dangi, and Oliver Steinbock. High-throughput robotic collection, imaging, and machine learning analysis of salt patterns: composition and concentration from dried droplet photos.Digital Discovery, 4(4):1030–1041, Feb 2025

  39. [47]

    Interaction networks for learning about objects, relations and physics.Advances in Neural Information Processing Systems, 29, Dec 2016

    Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics.Advances in Neural Information Processing Systems, 29, Dec 2016

  40. [48]

    Peerqa: A scientific question answering dataset from peer reviews

    Tim Baumgärtner, Ted Briscoe, and Iryna Gurevych. Peerqa: A scientific question answering dataset from peer reviews. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volum...

  41. [49]

    Toward machine learning optimization of experimental design.Nuclear Physics News, 31(1):25–28, Feb 2021

    Atılım Güneş Baydin, Kyle Cranmer, Pablo de Castro Manzano, Christophe Delaere, Denis Derkach, Julien Donini, Tommaso Dorigo, Andrea Giammanco, Jan Kieseler, Lukas Layer, et al. Toward machine learning optimization of experimental design.Nuclear Physics News, 31(1):25–28, Feb 2021

  42. [50]

    Unsupervisedpretrainingforfactverificationbylanguage model distillation.arXiv preprint arXiv:2309.16540, 2023

    AdriánBazaga, PietroLio, andGosMicklem. Unsupervisedpretrainingforfactverificationbylanguage model distillation.arXiv preprint arXiv:2309.16540, 2023. 52

  43. [51]

    Agentichypothesis: A survey on hypothesis generation using llm systems

    Adib Bazgir, Yuwen Zhang, et al. Agentichypothesis: A survey on hypothesis generation using llm systems. Towards Agentic AI for Science: Hypothesis Generation, Comprehension, Quantification, and Validation, Mar 2025

  44. [52]

    Paper recommender systems: a literature survey.International Journal on Digital Libraries, 17(4):305–338, Jul 2016

    Joeran Beel, Bela Gipp, Stefan Langer, and Corinna Breitinger. Paper recommender systems: a literature survey.International Journal on Digital Libraries, 17(4):305–338, Jul 2016

  45. [53]

    Joeran Beel, Min-Yen Kan, and Moritz Baumgart. Evaluating sakana’s ai scientist for autonomous research: Wishful thinking or an emerging reality towards’ artificial research intelligence’(ari)?arXiv preprint arXiv:2502.14297, 2025

  46. [54]

    Automatikz: Text-guided synthesis of scientific vector graphics with tikz.arXiv preprint arXiv:2310.00367, 2023

    Jonas Belouadi, Anne Lauscher, and Steffen Eger. Automatikz: Text-guided synthesis of scientific vector graphics with tikz.arXiv preprint arXiv:2310.00367, 2023

  47. [55]

    Tikzero: Zero-shot text-guided graphics program synthesis.arXiv preprint arXiv:2503.11509, 2025

    Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, and Simone Paolo Ponzetto. Tikzero: Zero-shot text-guided graphics program synthesis.arXiv preprint arXiv:2503.11509, 2025

  48. [56]

    SciBERT: A pretrained language model for scientific text

    Iz Beltagy, Kyle Lo, and Arman Cohan. SciBERT: A pretrained language model for scientific text. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors,Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Jo...

  49. [57]

    Mechanistic interpretability for ai safety–a review.arXiv preprint arXiv:2404.14082, 2024

    Leonard Bereska and Efstratios Gavves. Mechanistic interpretability for ai safety–a review.arXiv preprint arXiv:2404.14082, 2024

  50. [58]

    When automated assessment meets automated content generation: Examining text quality in the era of gpts.ACM Trans

    Marialena Bevilacqua, Kezia Oketch, Ruiyang Qin, Will Stamey, Xinyuan Zhang, Yi Gan, Kai Yang, and Ahmed Abbasi. When automated assessment meets automated content generation: Examining text quality in the era of gpts.ACM Trans. Inf. Syst., 43(2), January 2025. ISSN 1046-8188. ...

  51. [59]

    Peerassist: leveraging on paper-review interactions to predict peer review decisions

    Prabhat Kumar Bharti, Shashi Ranjan, Tirthankar Ghosal, Mayank Agrawal, and Asif Ekbal. Peerassist: leveraging on paper-review interactions to predict peer review decisions. InTowards Open and Trustworthy Digital Societies: 23rd International Conference on Asia-Pacific Digital...

  52. [60]

    Politepeer: does peer review hurt? a dataset to gauge politeness intensity in the peer reviews.Language Resources and Evaluation, 58(4):1291–1313, May 2024

    Prabhat Kumar Bharti, Meith Navlakha, Mayank Agarwal, and Asif Ekbal. Politepeer: does peer review hurt? a dataset to gauge politeness intensity in the peer reviews.Language Resources and Evaluation, 58(4):1291–1313, May 2024

  53. [61]

    General- purpose pre-trained large cellular models for single-cell transcriptomics.National Science Review, 11 (11):nwae340, Sep 2024

    Haiyang Bian, Yixin Chen, Erpai Luo, Xinze Wu, Minsheng Hao, Lei Wei, and Xuegong Zhang. General- purpose pre-trained large cellular models for single-cell transcriptomics.National Science Review, 11 (11):nwae340, Sep 2024

  54. [62]

    Helm: Highlighted evidence augmented language model for enhanced table-to-text generation.arXiv preprint arXiv:2311.08896, 2023

    Junyi Bian, Xiaolei Qin, Wuhe Zou, Mengzuo Huang, Congyi Luo, Ke Zhang, and Weidong Zhang. Helm: Highlighted evidence augmented language model for enhanced table-to-text generation.arXiv preprint arXiv:2311.08896, 2023. 53

  55. [63]

    Generating accurate and engaging research paper titles using nlp techniques

    Thulasi Bikku, Nirmala Rani Narimalla, Keerthi Konda, Anusha Nakkala, Avanti Yarlagadda, and B Sachuthananthan. Generating accurate and engaging research paper titles using nlp techniques. In International Conference on Innovations in Bio-Inspired Computing and Applications, p...

  56. [64]

    Using cognitive psychology to understand gpt-3

    Marcel Binz and Eric Schulz. Using cognitive psychology to understand gpt-3. Proceedings of the National Academy of Sciences, 120(6), February 2023. ISSN 1091-6490. doi: 10.1073/pnas. 2218523120. URL http://dx.doi.org/10.1073/pnas.2218523120

  57. [65]

    Designing collaborative intelligence systems for employee-ai service co-production.Journal of Service Research, page 10946705241238751, Mar 2024

    Marah Blaurock, Marion Büttgen, and Jeroen Schepers. Designing collaborative intelligence systems for employee-ai service co-production.Journal of Service Research, page 10946705241238751, Mar 2024

  58. [66]

    Improving gener- alization of robot locomotion policies via sharpness-aware reinforcement learning.arXiv preprint arXiv:2411.19732, 2024

    Severin Bochem, Eduardo Gonzalez-Sanchez, Yves Bicker, and Gabriele Fadini. Improving gener- alization of robot locomotion policies via sharpness-aware reinforcement learning.arXiv preprint arXiv:2411.19732, 2024

  59. [67]

    Colloquium: Machine learning in nuclear physics.Reviews of modern physics, 94(3):031003, Sep 2022

    Amber Boehnlein, Markus Diefenthaler, Nobuo Sato, Malachi Schram, Veronique Ziegler, Cristiano Fanelli, Morten Hjorth-Jensen, Tanja Horn, Michelle P Kuchera, Dean Lee, et al. Colloquium: Machine learning in nuclear physics.Reviews of modern physics, 94(3):031003, Sep 2022

  60. [68]

    Autonomous chemical research with large language models.Nature, 624(7992):570–578, Dec 2023

    Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models.Nature, 624(7992):570–578, Dec 2023

  61. [69]

    Artificial intelligence for literature reviews: Opportunities and challenges.Artificial Intelligence Review, 57(10):259, Aug 2024

    Francisco Bolanos, Angelo Salatino, Francesco Osborne, and Enrico Motta. Artificial intelligence for literature reviews: Opportunities and challenges.Artificial Intelligence Review, 57(10):259, Aug 2024

  62. [70]

    A non-factoid question-answering taxonomy

    Valeriia Bolotova, Vladislav Blinov, Falk Scholer, W Bruce Croft, and Mark Sanderson. A non-factoid question-answering taxonomy. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1196–1207, Jul 2022

  63. [71]

    Biomedlm: A 2.7 b parameter language model trained on biomedical text.arXiv preprint arXiv:2403.18421, 2024

    Elliot Bolton, Abhinav Venigalla, Michihiro Yasunaga, David Hall, Betty Xiong, Tony Lee, Roxana Daneshjou, Jonathan Frankle, Percy Liang, Michael Carbin, et al. Biomedlm: A 2.7 b parameter language model trained on biomedical text.arXiv preprint arXiv:2403.18421, 2024

  64. [72]

    Modest: A dataset for multi domain scientific title generation.Knowledge-Based Systems, page 113557, Jun 2025

    Necva Bölücü, Yunus Can Bilge, Dilber Çetintaş, and Zehra Yücel. Modest: A dataset for multi domain scientific title generation.Knowledge-Based Systems, page 113557, Jun 2025

  65. [73]

    Learning to split and rephrase from wikipedia edit history.arXiv preprint arXiv:1808.09468, 2018

    Jan A Botha, Manaal Faruqui, John Alex, Jason Baldridge, and Dipanjan Das. Learning to split and rephrase from wikipedia edit history.arXiv preprint arXiv:1808.09468, 2018

  66. [74]

    Kgvalidator: A framework for automatic validation of knowledge graph construction

    Jack Boylan, Shashank Mangla, Dominic Thorn, Demian Gholipour Ghalandari, Parsa Ghaffari, and Chris Hokamp. Kgvalidator: A framework for automatic validation of knowledge graph construction. arXiv preprint arXiv:2404.15923, 2024

  67. [75]

    Defame: Dynamic evidence- based fact-checking with multimodal experts.arXiv preprint arXiv:2412.10510, 2024

    Tobias Braun, Mark Rothermel, Marcus Rohrbach, and Anna Rohrbach. Defame: Dynamic evidence- based fact-checking with multimodal experts.arXiv preprint arXiv:2412.10510, 2024

  68. [76]

    Ai driven experiment calibration and control

    Thomas Britton, Cullan Bedwell, Abhijeet Chawhan, Julie Crowe, Naomi Jarvis, Torri Jeske, Nikhil Kalra, David Lawrence, and Diana McSpadden. Ai driven experiment calibration and control. InEPJ Web of Conferences, volume 295, page 02003. EDP Sciences, May 2024. 54

  69. [77]

    Generative artificial intelligence in anatomic pathology.Archives of Pathology & Laboratory Medicine, Apr 2025

    Victor Brodsky, Ehsan Ullah, Andrey Bychkov, Andrew H Song, Eric E Walk, Peter Louis, Ghulam Rasool, Rajendra S Singh, Faisal Mahmood, Marilyn M Bui, et al. Generative artificial intelligence in anatomic pathology.Archives of Pathology & Laboratory Medicine, Apr 2025

  70. [78]

    Closed-loop visuomotor control with generative expectation for robotic manipulation

    Qingwen Bu, Jia Zeng, Li Chen, Yanchao Yang, Guyue Zhou, Junchi Yan, Ping Luo, Heming Cui, Yi Ma, and Hongyang Li. Closed-loop visuomotor control with generative expectation for robotic manipulation. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  71. [79]

    Markus J Buehler. Accelerating scientific discovery with generative knowledge extraction, graph-based representation, and multimodal intelligent graph reasoning.Machine Learning: Science and Technology, 5(3):035083, Sep 2024

  72. [80]

    From large language models to multimodal ai: A scoping review on the potential of generative ai in medicine

    Lukas Buess, Matthias Keicher, Nassir Navab, Andreas Maier, and Soroosh Tayebi Arasteh. From large language models to multimodal ai: A scoping review on the potential of generative ai in medicine. arXiv preprint arXiv:2502.09242, 2025

  73. [81]

    How to build the virtual cell with artificial intelligence: Priorities and opportunities.Cell, 187(25):7045–7063, Dec 2024

    Charlotte Bunne, Yusuf Roohani, Yanay Rosen, Ankit Gupta, Xikun Zhang, Marcel Roed, Theo Alexandrov, Mohammed AlQuraishi, Patricia Brennan, Daniel B Burkhardt, et al. How to build the virtual cell with artificial intelligence: Priorities and opportunities.Cell, 187(25):7045–70...

  74. [82]

    Microvqa: A multimodal reasoning benchmark for microscopy-based scientific research

    JamesBurgess, JeffreyJNirschl, LauraBravo-Sánchez, AlejandroLozano, SanketRajanGupte, JesusG Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, et al. Microvqa: A multimodal reasoning benchmark for microscopy-based scientific research. InProceedings of the C...

  75. [83]

    Machine learning for molecular and materials science.Nature, 559(7715):547–555, Jul 2018

    Keith T Butler, Daniel W Davies, Hugh Cartwright, Olexandr Isayev, and Aron Walsh. Machine learning for molecular and materials science.Nature, 559(7715):547–555, Jul 2018

  76. [84]

    Modest: A dataset for multi domain scientific title generation.Knowledge-Based Systems, 321:113557, Jun 2025

    Necva Bölücü, Yunus Can Bilge, Dilber Çetintaş, and Zehra Yücel. Modest: A dataset for multi domain scientific title generation.Knowledge-Based Systems, 321:113557, Jun 2025. ISSN 0950-7051. doi: https://doi.org/10.1016/j.knosys.2025.113557. URLhttps://www.sciencedirect.com/ s...

  77. [85]

    On gradient-like explanation under a black-box setting: when black-box explanations become as good as white-box.arXiv preprint arXiv:2308.09381, 2023

    Yi Cai and Gerhard Wunder. On gradient-like explanation under a black-box setting: when black-box explanations become as good as white-box.arXiv preprint arXiv:2308.09381, 2023

  78. [86]

    Reyes Calderon and Francisco Herrera. And plato met chatgpt: an ethical reflection on the use of chatbots in scientific research writing, with a particular focus on the social sciences.Humanities and Social Sciences Communications, 12(1):1–13, May 2025

  79. [87]

    Science acceleration and accessibility with self-driving labs.Nature Communications, 16(1):3856, Apr 2025

    Richard B Canty, Jeffrey A Bennett, Keith A Brown, Tonio Buonassisi, Sergei V Kalinin, John R Kitchin, Benji Maruyama, Robert G Moore, Joshua Schrier, Martin Seifrid, et al. Science acceleration and accessibility with self-driving labs.Nature Communications, 16(1):3856, Apr 2025

  80. [88]

    Tablemaster: A recipe to advance table understanding with language models

    Lang Cao and Hanbing Liu. Tablemaster: A recipe to advance table understanding with language models. arXiv preprint arXiv:2501.19378, 2025

  81. [89]

    Agents for self-driving laboratories applied to quantum computing

    Shuxiang Cao, Zijian Zhang, Mohammed Alghadeer, Simone D Fasciati, Michele Piscitelli, Mustafa Bakr, Peter Leek, and Alán Aspuru-Guzik. Agents for self-driving laboratories applied to quantum computing. arXiv preprint arXiv:2412.07978, 2024. 55

  82. [90]

    Figuring out figures: Using textual references to caption scientific figures

    Stanley Cao and Kevin Liu. Figuring out figures: Using textual references to caption scientific figures. arXiv preprint arXiv:2407.11008, 2024

  83. [91]

    Can large language models detect misinformation in scientific news reporting?arXiv preprint arXiv:2402.14268, 2024

    Yupeng Cao, Aishwarya Muralidharan Nair, Elyon Eyimife, Nastaran Jamalipour Soofi, KP Subbal- akshmi, John R Wullert II, Chumki Basu, and David Shallcross. Can large language models detect misinformation in scientific news reporting?arXiv preprint arXiv:2402.14268, 2024

  84. [92]

    Citebart: Learning to generate citations for local citation recommen- dation

    Ege Yiğit Çelik and Selma Tekir. Citebart: Learning to generate citations for local citation recommen- dation. arXiv preprint arXiv:2412.17534, 2024

  85. [93]

    Art or artifice? large language models and the false promise of creativity

    Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. Art or artifice? large language models and the false promise of creativity. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pages 1–34, May 2024

  86. [94]

    Automated focused feedback generation for scientific writing assistance

    Eric Chamoun, Michael Schlichtkrull, and Andreas Vlachos. Automated focused feedback generation for scientific writing assistance. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Findings of the Association for Computational Linguistics: ACL 2024, pages 9742–9763, Ba...

  87. [95]

    Mle-bench: Evaluating machine learning agents on machine learning engineering.arXiv preprint arXiv:2410.07095, 2024

    Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al. Mle-bench: Evaluating machine learning agents on machine learning engineering.arXiv preprint arXiv:2410.07095, 2024

  88. [96]

    From lived experi- ence to insight: Unpacking the psychological risks of using ai conversational agents.arXiv preprint arXiv:2412.07951, 2024

    Mohit Chandra, Suchismita Naik, Denae Ford, Ebele Okoli, Munmun De Choudhury, Mahsa Ershadi, Gonzalo Ramos, Javier Hernandez, Ananya Bhattacharjee, Shahed Warreth, et al. From lived experi- ence to insight: Unpacking the psychological risks of using ai conversational agents.ar...

  89. [97]

    Sridbench: Benchmark of scientific research illustration drawing of image generation model.arXiv preprint arXiv:2505.22126, 2025

    Yifan Chang, Yukang Feng, Jianwen Sun, Jiaxin Ai, Chuanhao Li, S Kevin Zhou, and Kaipeng Zhang. Sridbench: Benchmark of scientific research illustration drawing of image generation model.arXiv preprint arXiv:2505.22126, 2025

  90. [98]

    Treereview: A dynamic tree of questions framework for deep and efficient llm-based scientific peer review

    Yuan Chang, Ziyue Li, Hengyuan Zhang, Yuanbo Kong, Yanru Wu, Zhijiang Guo, and Ngai Wong. Treereview: A dynamic tree of questions framework for deep and efficient llm-based scientific peer review. arXiv preprint arXiv:2506.07642, 2025

  91. [99]

    Thetorontopapermatchingsystem: anautomatedpaper-reviewer assignment system

    LaurentCharlinandRichardZemel. Thetorontopapermatchingsystem: anautomatedpaper-reviewer assignment system. May 2013

  92. [100]

    A framework for optimizing paper matching

    Laurent Charlin, Richard Zemel, and Craig Boutilier. A framework for optimizing paper matching. In Proceedings of the Twenty-Seventh Conference on Uncertainty in Artificial Intelligence, pages 86–95, Jul 2011

  93. [101]

    Llm based exploratory data analysis using bigquery data canvas, Oct 2024

    Rajdip Chaudhuri. Llm based exploratory data analysis using bigquery data canvas, Oct 2024. URL https://medium.com/google-cloud/ llm-based-exploratory-data-analysis-using-bigquery-data-canvas-42fbecb9f009 . LLM Based Exploratory Data Analysis Using BigQuery Data Canvas

  94. [102]

    Graph networks as a universal machine learning framework for molecules and crystals.Chemistry of Materials, 31(9):3564–3572, Apr 2019

    Chi Chen, Weike Ye, Yunxing Zuo, Chen Zheng, and Shyue Ping Ong. Graph networks as a universal machine learning framework for molecules and crystals.Chemistry of Materials, 31(9):3564–3572, Apr 2019. 56

  95. [103]

    Mlr-bench: Evaluating ai agents on open-ended machine learning research.arXiv preprint arXiv:2505.19955, 2025

    Hui Chen, Miao Xiong, Yujie Lu, Wei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, and Bryan Hooi. Mlr-bench: Evaluating ai agents on open-ended machine learning research.arXiv preprint arXiv:2505.19955, 2025

  96. [104]

    Automatic generation of related work through summarizing citations

    Jingqiang Chen and Hai Zhuge. Automatic generation of related work through summarizing citations. Concurrency and Computation: Practice and Experience, 31(3):e4261, Sep 2019

  97. [105]

    Structuring scientific innovation: A framework for modeling and discovering impactful knowledge combinations

    Junlan Chen, Kexin Zhang, Daifeng Li, Yangyang Feng, Yuxuan Zhang, and Bowen Deng. Structuring scientific innovation: A framework for modeling and discovering impactful knowledge combinations. arXiv preprint arXiv:2503.18865, 2025

  98. [106]

    Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021

  99. [107]

    Xtragpt: Llms for human-ai collaboration on controllable academic paper revision

    Nuo Chen, Andre Lin HuiKai, Jiaying Wu, Junyi Hou, Zining Zhang, Qian Wang, Xidong Wang, and Bingsheng He. Xtragpt: Llms for human-ai collaboration on controllable academic paper revision. arXiv preprint arXiv:2505.11336, 2025

  100. [108]

    Unlock- ing the capabilities of thought: A reasoning boundary framework to quantify and opti- mize chain-of-thought

    Qiguang Chen, Libo Qin, Jiaqi Wang, Jingxuan Zhou, and Wanxiang Che. Unlock- ing the capabilities of thought: A reasoning boundary framework to quantify and opti- mize chain-of-thought. Advances in Neural Information Processing Systems, 37:54872–54904, Sep 2024. URL https://pr...

  101. [109]

    M3CoT:Anovelbenchmark for multi-domain multi-step multi-modal chain-of-thought

    QiguangChen, LiboQin, JinZhang, ZhiChen, XiaoXu, andWanxiangChe. M3CoT:Anovelbenchmark for multi-domain multi-step multi-modal chain-of-thought. pages 8199–8221, August 2024. doi: 10.18653/v1/2024.acl-long.446. URL https://aclanthology.org/2024.acl-long.446/

  102. [110]

    Rbf++: Quantifying and optimizing reasoning boundaries across measurable and unmeasurable capabilities for chain-of-thought reasoning.arXiv preprint arXiv:2505.13307, 2025

    Qiguang Chen, Libo Qin, Jinhao Liu, Yue Liao, Jiaqi Wang, Jingxuan Zhou, and Wanxiang Che. Rbf++: Quantifying and optimizing reasoning boundaries across measurable and unmeasurable capabilities for chain-of-thought reasoning.arXiv preprint arXiv:2505.13307, 2025

  103. [111]

    Towards reasoning era: A survey of long chain-of-thought for reasoning large language models.arXiv preprint arXiv:2503.09567, 2025

    Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiannan Guan, Peng Wang, Mengkang Hu, Yuhang Zhou, Te Gao, and Wangxiang Che. Towards reasoning era: A survey of long chain-of-thought for reasoning large language models.arXiv preprint arXiv:2503.09567, 2025

  104. [112]

    Ecm: A unified electronic circuit model for explaining the emergence of in-context learning and chain-of-thought in large language model.arXiv preprint arXiv:2502.03325, 2025

    Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiaqi Wang, Mengkang Hu, Zhi Chen, Wanxiang Che, and Ting Liu. Ecm: A unified electronic circuit model for explaining the emergence of in-context learning and chain-of-thought in large language model.arXiv preprint arXiv:2502.0...

  105. [113]

    Ai-driven automation can become the foundation of next-era science of science research.arXiv preprint arXiv:2505.12039, 2025

    Renqi Chen, Haoyang Su, Shixiang Tang, Zhenfei Yin, Qi Wu, Hui Li, Ye Sun, Nanqing Dong, Wanli Ouyang, and Philip Torr. Ai-driven automation can become the foundation of next-era science of science research.arXiv preprint arXiv:2505.12039, 2025

  106. [114]

    Bridging social psychology and llm reasoning: Conflict-aware meta-review generation via cognitive alignment.arXiv preprint arXiv:2503.13879, 2025

    Wei Chen, Han Ding, Meng Yuan, Zhao Zhang, Deqing Wang, and Fuzhen Zhuang. Bridging social psychology and llm reasoning: Conflict-aware meta-review generation via cognitive alignment.arXiv preprint arXiv:2503.13879, 2025

  107. [115]

    Theoremqa: A theorem-driven question answering dataset.arXiv preprint arXiv:2305.12524, 2023

    Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia. Theoremqa: A theorem-driven question answering dataset.arXiv preprint arXiv:2305.12524, 2023. 57

  108. [116]

    Nano & ai: A nobel partnership, Nov 2024

    Xiaodong Chen, Jillian M Buriak, Mathieu Salanne, and Huolin Xin. Nano & ai: A nobel partnership, Nov 2024

  109. [117]

    Teaching large language models to self-debug

    Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128, 2023

  110. [118]

    Capturing relations between scientific papers: An abstractive model for related work section generation

    Xiuying Chen, Hind Alamro, Mingzhe Li, Shen Gao, Xiangliang Zhang, Dongyan Zhao, and Rui Yan. Capturing relations between scientific papers: An abstractive model for related work section generation. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors,Proceeding...

  111. [119]

    Target-aware abstractive related work generation with contrastive learning

    Xiuying Chen, Hind Alamro, Mingzhe Li, Shen Gao, Rui Yan, Xin Gao, and Xiangliang Zhang. Target-aware abstractive related work generation with contrastive learning. InProceedings of the 45th international ACM SIGIR conference on research and development in information retrieva...

  112. [120]

    Scholarchemqa: Unveiling the power of language models in chemical research question answering.arXiv preprint arXiv:2407.16931, 2024

    Xiuying Chen, Tairan Wang, Taicheng Guo, Kehan Guo, Juexiao Zhou, Haoyang Li, Mingchen Zhuge, Jürgen Schmidhuber, Xin Gao, and Xiangliang Zhang. Scholarchemqa: Unveiling the power of language models in chemical research question answering.arXiv preprint arXiv:2407.16931, 2024

  113. [121]

    Predicting field experiments with large language models

    Yaoyu Chen, Yuheng Hu, and Yingda Lu. Predicting field experiments with large language models. arXiv preprint arXiv:2504.01167, 2025

  114. [122]

    Uniter: Universal image-text representation learning

    Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. InEuropean conference on computer vision, pages 104–120. Springer, Sep 2020

  115. [123]

    Reinforcing clinical decision support through multi-agent systems and ethical ai governance.arXiv preprint arXiv:2504.03699, 2025

    Ying-Jung Chen, Ahmad Albarqawi, and Chi-Sheng Chen. Reinforcing clinical decision support through multi-agent systems and ethical ai governance.arXiv preprint arXiv:2504.03699, 2025

  116. [124]

    Genept: a simple but effective foundation model for genes and cells built from chatgpt.bioRxiv, pages 2023–10, Mar 2024

    Yiqun Chen and James Zou. Genept: a simple but effective foundation model for genes and cells built from chatgpt.bioRxiv, pages 2023–10, Mar 2024

  117. [125]

    The emergence of economic rationality of gpt

    Yiting Chen, Tracy Xiao Liu, You Shan, and Songfa Zhong. The emergence of economic rationality of gpt. Proceedings of the National Academy of Sciences, 120(51):e2316205120, Dec 2023

  118. [126]

    What are the essential factors in crafting effective long context multi-hop instruction datasets? insights and best practices.arXiv preprint arXiv:2409.01893, 2024

    Zhi Chen, Qiguang Chen, Libo Qin, Qipeng Guo, Haijun Lv, Yicheng Zou, Wanxiang Che, Hang Yan, Kai Chen, and Dahua Lin. What are the essential factors in crafting effective long context multi-hop instruction datasets? insights and best practices.arXiv preprint arXiv:2409.01893, 2024

  119. [127]

    Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery.arXiv preprint arXiv:2410.05080, 2024

    Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, et al. Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery.arXiv preprint arXiv:2410.05080, 2024

  120. [128]

    Artificial intelligence-assisted academic writing: recommendations for ethical use.Advances in Simulation, 10(1):22, Apr 2025

    Adam Cheng, Aaron Calhoun, and Gabriel Reedy. Artificial intelligence-assisted academic writing: recommendations for ethical use.Advances in Simulation, 10(1):22, Apr 2025

  121. [129]

    Language modeling by language models.arXiv preprint arXiv:2506.20249, 2025

    Junyan Cheng, Peter Clark, and Kyle Richardson. Language modeling by language models.arXiv preprint arXiv:2506.20249, 2025. 58

  122. [130]

    Causal learning for socially responsible ai.arXiv preprint arXiv:2104.12278, 2021

    Lu Cheng, Ahmadreza Mosallanezhad, Paras Sheth, and Huan Liu. Causal learning for socially responsible ai.arXiv preprint arXiv:2104.12278, 2021

  123. [131]

    Biasfilter: An inference-time debiasing framework for large language models.arXiv preprint arXiv:2505.23829, 2025

    Xiaoqing Cheng, Ruizhe Chen, Hongying Zan, Yuxiang Jia, and Min Peng. Biasfilter: An inference-time debiasing framework for large language models.arXiv preprint arXiv:2505.23829, 2025

  124. [132]

    Ai-generated literature reviews threaten scientific progress.Nature, 641(8064):852–852, 2025

    Xusen Cheng and Lulu Zhang. Ai-generated literature reviews threaten scientific progress.Nature, 641(8064):852–852, 2025

  125. [133]

    Chartreader: A unified framework for chart derendering and comprehension without heuristic rules

    Zhi-Qi Cheng, Qi Dai, and Alexander G Hauptmann. Chartreader: A unified framework for chart derendering and comprehension without heuristic rules. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 22202–22213, Apr 2023

  126. [134]

    Visual thoughts: A unified perspective of understanding multimodal chain-of-thought.arXiv preprint arXiv:2505.15510, 2025

    Zihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang, Weiyun Wang, Hao Fei, Yidong Wang, Alex Jinpeng Wang, Zhi Chen, Wanxiang Che, et al. Visual thoughts: A unified perspective of understanding multimodal chain-of-thought.arXiv preprint arXiv:2505.15510, 2025

  127. [135]

    Comt: A novel benchmark for chain of multi-modal thought on large vision-language models

    Zihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei, Xiaocheng Feng, Wanxiang Che, Min Li, and Libo Qin. Comt: A novel benchmark for chain of multi-modal thought on large vision-language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 23678...

  128. [136]

    Accelerating materials language processing with large language models

    Jaewoong Choi and Byungju Lee. Accelerating materials language processing with large language models. Communications Materials, 5(1):13, Feb 2024

  129. [137]

    Self-critique guided iterative reasoning for multi-hop question answering.arXiv preprint arXiv:2505.19112, 2025

    Zheng Chu, Huiming Fan, Jingchang Chen, Qianyu Wang, Mingda Yang, Jiafeng Liang, Zhongjie Wang, Hao Li, Guo Tang, Ming Liu, et al. Self-critique guided iterative reasoning for multi-hop question answering.arXiv preprint arXiv:2505.19112, 2025

  130. [138]

    Automatic large language model evaluation via peer review

    Zhumin Chu, Qingyao Ai, Yiteng Tu, Haitao Li, and Yiqun Liu. Automatic large language model evaluation via peer review. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management, pages 384–393, Oct 2024

  131. [139]

    Pre: A peer review based large language model evaluator.arXiv preprint arXiv:2401.15641, 2024

    Zhumin Chu, Qingyao Ai, Yiteng Tu, Haitao Li, and Yiqun Liu. Pre: A peer review based large language model evaluator.arXiv preprint arXiv:2401.15641, 2024

  132. [140]

    Crosslingual capabilities and knowledge barriers in multilingual large language models.arXiv preprint arXiv:2406.16135, 2024

    Lynn Chua, Badih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar, Pasin Manurangsi, Amer Sinha, Chulin Xie, and Chiyuan Zhang. Crosslingual capabilities and knowledge barriers in multilingual large language models.arXiv preprint arXiv:2406.16135, 2024

  133. [141]

    Ai-experimentsineducation: Anai-drivenrandomized controlled trial for higher education research.Education and Information Technologies, 29(15):19649– 19677, 2024

    IlkerCingillioglu,UriGal,andArtemProkhorov. Ai-experimentsineducation: Anai-drivenrandomized controlled trial for higher education research.Education and Information Technologies, 29(15):19649– 19677, 2024

  134. [142]

    How well do large language models understand tables in materials science?Integrating Materials and Manufac- turing Innovation, 13(3):669–687, Jul 2024

    Defne Circi, Ghazal Khalighinejad, Anlan Chen, Bhuwan Dhingra, and L Catherine Brinson. How well do large language models understand tables in materials science?Integrating Materials and Manufac- turing Innovation, 13(3):669–687, Jul 2024. URLhttps://link.springer.com/article/...

  135. [143]

    Artificial intelligence in cancer research: learning at different levels of data granularity.Molecular oncology, 15(4):817–829, Apr 2021

    Davide Cirillo, Iker Núñez-Carpintero, and Alfonso Valencia. Artificial intelligence in cancer research: learning at different levels of data granularity.Molecular oncology, 15(4):817–829, Apr 2021. 59

  136. [144]

    Boolq: Exploring the surprising difficulty of natural yes/no questions.arXiv preprint arXiv:1905.10044, 2019

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions.arXiv preprint arXiv:1905.10044, 2019

  137. [145]

    Wordcraft: A human-ai collaborative editor for story writing.arXiv preprint arXiv:2107.07430, 2021

    Andy Coenen, Luke Davis, Daphne Ippolito, Emily Reif, and Ann Yuan. Wordcraft: A human-ai collaborative editor for story writing.arXiv preprint arXiv:2107.07430, 2021

  138. [146]

    Large language models for causal hypothesis generation in science.Machine Learning: Science and Technology, 6(1):013001, Jan 2025

    Kai-Hendrik Cohrs, Emiliano Diaz, Vasileios Sitokonstantinou, Gherardo Varando, and Gustau Camps- Valls. Large language models for causal hypothesis generation in science.Machine Learning: Science and Technology, 6(1):013001, Jan 2025

  139. [147]

    Unsupervised cross- lingual representation learning at scale.arXiv preprint arXiv:1911.02116, 2019

    AlexisConneau, KartikayKhandelwal, NamanGoyal, VishravChaudhary, GuillaumeWenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross- lingual representation learning at scale.arXiv preprint arXiv:1911.02116, 2019

  140. [148]

    Transformers and genome language models

    Micaela E Consens, Cameron Dufault, Michael Wainberg, Duncan Forster, Mehran Karimzadeh, Hani Goodarzi, Fabian J Theis, Alan Moses, and Bo Wang. Transformers and genome language models. Nature Machine Intelligence, pages 1–17, Mar 2025

  141. [149]

    Relevai-reviewer: A benchmark on ai reviewers for survey paper relevance.arXiv preprint arXiv:2406.10294, 2024

    Paulo Henrique Couto, Quang Phuoc Ho, Nageeta Kumari, Benedictus Kent Rachmat, Thanh Gia Hieu Khuong, Ihsan Ullah, and Lisheng Sun-Hosoya. Relevai-reviewer: A benchmark on ai reviewers for survey paper relevance.arXiv preprint arXiv:2406.10294, 2024

  142. [150]

    A human-LLM note-taking system with case-based reasoning as framework for scientific discovery

    Douglas B Craig. A human-LLM note-taking system with case-based reasoning as framework for scientific discovery. In Peter Jansen, Bhavana Dalvi Mishra, Harsh Trivedi, Bodhisattwa Prasad Ma- jumder, Tom Hope, Tushar Khot, Doug Downey, and Eric Horvitz, editors,Proceedings of th...

  143. [151]

    Lagrangian neural networks.arXiv preprint arXiv:2003.04630, 2020

    Miles Cranmer, Sam Greydanus, Stephan Hoyer, Peter Battaglia, David Spergel, and Shirley Ho. Lagrangian neural networks.arXiv preprint arXiv:2003.04630, 2020

  144. [152]

    scgpt: toward building a foundation model for single-cell multi-omics using generative ai.Nature Methods, 21(8):1470–1480, Feb 2024

    Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengning Luo, Nan Duan, and Bo Wang. scgpt: toward building a foundation model for single-cell multi-omics using generative ai.Nature Methods, 21(8):1470–1480, Feb 2024

  145. [153]

    Lumi-lab: a foundation model-driven autonomous platform enabling discovery of new ionizable lipid designs for mrna delivery.BioRxiv, pages 2025–02, Feb 2025

    Haotian Cui, Yue Xu, Kuan Pang, Gen Li, Fanglin Gong, Bo Wang, and Bowen Li. Lumi-lab: a foundation model-driven autonomous platform enabling discovery of new ionizable lipid designs for mrna delivery.BioRxiv, pages 2025–02, Feb 2025

  146. [154]

    Can ai replace human subjects? a large-scale replication of psychological experiments with llms.arXiv preprint arXiv:2409.00128, 2024

    Ziyan Cui, Ning Li, and Huaikang Zhou. Can ai replace human subjects? a large-scale replication of psychological experiments with llms.arXiv preprint arXiv:2409.00128, 2024

  147. [155]

    Autonomous mobile robots for exploratory synthetic chemistry.Nature, pages 1–8, Nov 2024

    Tianwei Dai, Sriram Vijayakrishnan, Filip T Szczypiński, Jean-François Ayme, Ehsan Simaei, Thomas Fellowes, Rob Clowes, Lyubomir Kotopanov, Caitlin E Shields, Zhengxue Zhou, et al. Autonomous mobile robots for exploratory synthetic chemistry.Nature, pages 1–8, Nov 2024

  148. [156]

    Adaptive ai decision interface for autonomous electronic material discovery.arXiv preprint arXiv:2504.13344, 2025

    Yahao Dai, Henry Chan, Aikaterini Vriza, Fredrick Kim, Yunfei Wang, Wei Liu, Naisong Shan, Jing Xu, Max Weires, Yukun Wu, et al. Adaptive ai decision interface for autonomous electronic material discovery.arXiv preprint arXiv:2504.13344, 2025. 60

  149. [157]

    Claimver: Explainable claim-level verification and evidence attribution of text through knowledge graphs.arXiv preprint arXiv:2403.09724, 2024

    Preetam Prabhu Srikar Dammu, Himanshu Naidu, Mouly Dewan, YoungMin Kim, Tanya Roosta, Aman Chadha, and Chirag Shah. Claimver: Explainable claim-level verification and evidence attribution of text through knowledge graphs.arXiv preprint arXiv:2403.09724, 2024

  150. [158]

    Machine learning-aided inverse design and discovery of novel polymeric materials for membrane separation

    Raghav Dangayach, Nohyeong Jeong, Elif Demirel, Nigmet Uzal, Victor Fung, and Yongsheng Chen. Machine learning-aided inverse design and discovery of novel polymeric materials for membrane separation. Environmental Science & Technology, 59(2):993–1012, Dec 2024

  151. [159]

    Marg: Multi-agent review generation for scientific papers.arXiv preprint arXiv:2401.04259, 2024

    Mike D’Arcy, Tom Hope, Larry Birnbaum, and Doug Downey. Marg: Multi-agent review generation for scientific papers.arXiv preprint arXiv:2401.04259, 2024

  152. [160]

    Glimpse: Pragmatically informative multi-document summarization for scholarly reviews.arXiv preprint arXiv:2406.07359, 2024

    Maxime Darrin, Ines Arous, Pablo Piantanida, and Jackie CK Cheung. Glimpse: Pragmatically informative multi-document summarization for scholarly reviews.arXiv preprint arXiv:2406.07359, 2024

  153. [161]

    Organa: a robotic assistant for automated chemistry experimentation and characterization.Matter, 8(2), Feb 2025

    Kourosh Darvish, Marta Skreta, Yuchi Zhao, Naruki Yoshikawa, Sagnik Som, Miroslav Bogdanovic, Yang Cao, Han Hao, Haoping Xu, Alán Aspuru-Guzik, et al. Organa: a robotic assistant for automated chemistry experimentation and characterization.Matter, 8(2), Feb 2025

  154. [162]

    The state of human-centered nlp technology for fact-checking.Information processing & management, 60(2):103219, Mar 2023

    Anubrata Das, Houjiang Liu, Venelin Kovatchev, and Matthew Lease. The state of human-centered nlp technology for fact-checking.Information processing & management, 60(2):103219, Mar 2023

  155. [163]

    Empowering ai as autonomous researchers: Evaluating llms in generating novel research ideas through automated metrics

    Debajyoti Dasgupta, Arijit Mondal, and Partha Pratim Chakrabarti. Empowering ai as autonomous researchers: Evaluating llms in generating novel research ideas through automated metrics. In2nd AI4Research Workshop: Towards a Knowledge-grounded Scientific Research Lifecycle, Dec 2024

  156. [164]

    End-to- end differentiable physics for learning and control.Advances in Neural Information Processing Systems, 31, Dec 2018

    Filipe de Avila Belbute-Peres, Kevin Smith, Kelsey Allen, Josh Tenenbaum, and J Zico Kolter. End-to- end differentiable physics for learning and control.Advances in Neural Information Processing Systems, 31, Dec 2018

  157. [165]

    Molgan: An implicit generative model for small molecular graphs

    Nicola De Cao and Thomas Kipf. Molgan: An implicit generative model for small molecular graphs. arXiv preprint arXiv:1805.11973, 2018

  158. [166]

    Chatgpt for textual analysis? how to use generative llms in accounting research

    Ties de Kok. Chatgpt for textual analysis? how to use generative llms in accounting research. Management Science, Jan 2025. doi: 10.1287/mnsc.2023.03253. URL https://doi.org/10. 1287/mnsc.2023.03253. Published online in Articles in Advance, 13 Jan 2025

  159. [167]

    Deiner, Vlad Honcharov, Jiawei Li, Tim K

    Michael S. Deiner, Vlad Honcharov, Jiawei Li, Tim K. Mackey, Travis C. Porco, and Urmimala Sarkar. Large language models can enable inductive thematic analysis of a social media corpus in a single prompt: Human validation study.JMIR Infodemiology, 4:e59641, August 2024. doi: 1...

  160. [168]

    Automatic related work section generation by sentence extraction and reordering

    Zekun Deng, Zixin Zeng, Weiye Gu, Jiawen Ji, and Bolin Hua. Automatic related work section generation by sentence extraction and reordering. InAII@ iConference, pages 101–110, Jan 2021

  161. [169]

    Autoscilab: A self-driving laboratory for interpretable scientific discovery

    Saaketh Desai, Sadhvikas Addamane, Jeffrey Y Tsao, Igal Brener, Laura P Swiler, Remi Dingreville, and Prasad P Iyer. Autoscilab: A self-driving laboratory for interpretable scientific discovery. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages ...

  162. [170]

    Ms2: Multi-document summarization of medical studies.arXiv preprint arXiv:2104.06486, 2021

    JayDeYoung, IzBeltagy, MadeleinevanZuylen, BaileyKuehl, andLucyLuWang. Ms2: Multi-document summarization of medical studies.arXiv preprint arXiv:2104.06486, 2021. 61

  163. [171]

    Streamlining the review process: Ai-generated annotations in research manuscripts.arXiv preprint arXiv:2412.00281, 2024

    Oscar Díaz, Xabier Garmendia, and Juanan Pereira. Streamlining the review process: Ai-generated annotations in research manuscripts.arXiv preprint arXiv:2412.00281, 2024

  164. [172]

    Can ai language models replace hu- man participants? Trends in Cognitive Sciences, 27(7):597–600, Jul 2023

    Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. Can ai language models replace hu- man participants? Trends in Cognitive Sciences, 27(7):597–600, Jul 2023. ISSN 1364-6613. doi: https://doi.org/10.1016/j.tics.2023.04.008. URL https://www.sciencedirect.com/ science/artic...

  165. [173]

    Automating exploratory proteomics research via language models.arXiv preprint arXiv:2411.03743, 2024

    Ning Ding, Shang Qu, Linhai Xie, Yifei Li, Zaoqu Liu, Kaiyan Zhang, Yibai Xiong, Yuxin Zuo, Zhangren Chen, Ermo Hua, et al. Automating exploratory proteomics research via language models.arXiv preprint arXiv:2411.03743, 2024

  166. [174]

    Popular and/or prestigious? measures of scholarly esteem.Information processing & management, 47(1):80–96, Jan 2011

    Ying Ding and Blaise Cronin. Popular and/or prestigious? measures of scholarly esteem.Information processing & management, 47(1):80–96, Jan 2011

  167. [175]

    Oarelatedwork: A large-scale dataset of related work sections with full-texts from open access sources.arXiv preprint arXiv:2405.01930, 2024

    Martin Docekal, Martin Fajcik, and Pavel Smrz. Oarelatedwork: A large-scale dataset of related work sections with full-texts from open access sources.arXiv preprint arXiv:2405.01930, 2024

  168. [176]

    Aide: Human-level performance on data science competitions, Apr 2023

    Schmidt Dominik, Jiang Zhengyao, and Wu Yuxiang. Aide: Human-level performance on data science competitions, Apr 2023. URLhttps://www.weco.ai/blog/technical-report. AIDE

  169. [177]

    Toward the end-to-end optimization of particle physics instruments with differentiable programming.Reviews in Physics, 10: 100085, Jun 2023

    Tommaso Dorigo, Andrea Giammanco, Pietro Vischia, Max Aehle, Mateusz Bawaj, Alexey Boldyrev, Pablo de Castro Manzano, Denis Derkach, Julien Donini, Auralee Edelen, et al. Toward the end-to-end optimization of particle physics instruments with differentiable programming.Reviews...

  170. [178]

    Artificial intelligence in peer review: enhancing efficiency while preserving integrity.Journal of Korean medical science, 40(7), Feb 2025

    Bohdana Doskaliuk, Olena Zimba, Marlen Yessirkepov, Iryna Klishch, and Roman Yatsyshyn. Artificial intelligence in peer review: enhancing efficiency while preserving integrity.Journal of Korean medical science, 40(7), Feb 2025

  171. [179]

    Semi-supervised classification with novelty detection using support vector machines and linear discriminant analysis

    Ryan S Dove, Roy J Hartfield, and Mark Carpenter. Semi-supervised classification with novelty detection using support vector machines and linear discriminant analysis. InAIAA SCITECH 2025 Forum, page 0705, Jan 2025

  172. [180]

    Yu, and Wenpeng Yin

    Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Ji...

  173. [181]

    Deepresearch bench: A comprehensive benchmark for deep research agents.arXiv preprint arXiv:2506.11763, 2025

    Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. Deepresearch bench: A comprehensive benchmark for deep research agents.arXiv preprint arXiv:2506.11763, 2025

  174. [182]

    Read, revise, repeat: A system demonstration for human-in-the-loop iterative text revision.arXiv preprint arXiv:2204.03685, 2022

    Wanyu Du, Zae Myung Kim, Vipul Raheja, Dhruv Kumar, and Dongyeop Kang. Read, revise, repeat: A system demonstration for human-in-the-loop iterative text revision.arXiv preprint arXiv:2204.03685, 2022. 62

  175. [183]

    NLPeer: A unified resource for the computational study of peer review

    Nils Dycke, Ilia Kuznetsov, and Iryna Gurevych. NLPeer: A unified resource for the computational study of peer review. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volum...

  176. [184]

    Axolotl: fairness through assisted self-debiasing of large language model outputs.arXiv preprint arXiv:2403.00198, 2024

    Sana Ebrahimi, Kaiwen Chen, Abolfazl Asudeh, Gautam Das, and Nick Koudas. Axolotl: fairness through assisted self-debiasing of large language model outputs.arXiv preprint arXiv:2403.00198, 2024

  177. [185]

    Brice Edelman and Jeffrey Skolnick. Valsci: an open-source, self-hostable literature review utility for automated large-batch scientific claim verification using large language models.BMC bioinformatics, 26(1):1–25, May 2025

  178. [186]

    A data science roadmap for open science organizations engaged in early-stage drug discovery.Nature Communications, 15(1): 5640, Jul 2024

    Kristina Edfeldt, Aled M Edwards, Ola Engkvist, Judith Günther, Matthew Hartley, David G Hulcoop, Andrew R Leach, Brian D Marsden, Amelie Menge, Leonie Misquitta, et al. A data science roadmap for open science organizations engaged in early-stage drug discovery.Nature Communic...

  179. [187]

    Transforming science with large lan- guage models: A survey on ai-assisted scientific discovery, experimentation, content generation, and evaluation

    Steffen Eger, Yong Cao, Jennifer D’Souza, Andreas Geiger, Christian Greisinger, Stephanie Gross, Yufang Hou, Brigitte Krenn, Anne Lauscher, Yizhi Li, et al. Transforming science with large lan- guage models: A survey on ai-assisted scientific discovery, experimentation, conten...

  180. [188]

    Explainable ai reloaded: Challenging the xai status quo in the era of large language models

    Upol Ehsan and Mark Riedl. Explainable ai reloaded: Challenging the xai status quo in the era of large language models. InProceedings of the Halfway to the Future Symposium, pages 1–8, Oct 2024

  181. [189]

    Accelerating the discovery of abiotic vesicles with ai-guided automated experimentation.Langmuir, 41(1):858–867, 2024

    Christelle Ekosso, Hao Liu, Avery Glagovich, Dustin Nguyen, Sarah Maurer, and Joshua Schrier. Accelerating the discovery of abiotic vesicles with ai-guided automated experimentation.Langmuir, 41(1):858–867, 2024

  182. [190]

    Automated justification production for claim veracity in fact checking: A survey on architectures and approaches.arXiv preprint arXiv:2407.12853, 2024

    Islam Eldifrawi, Shengrui Wang, and Amine Trabelsi. Automated justification production for claim veracity in fact checking: A survey on architectures and approaches.arXiv preprint arXiv:2407.12853, 2024

  183. [191]

    Text editing by command.arXiv preprint arXiv:2010.12826, 2020

    Felix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao, and Bill Dolan. Text editing by command.arXiv preprint arXiv:2010.12826, 2020

  184. [192]

    Generating full length wikipedia biographies: The impact of gender bias on the retrieval-based generation of women biographies.arXiv preprint arXiv:2204.05879, 2022

    Angela Fan and Claire Gardent. Generating full length wikipedia biographies: The impact of gender bias on the retrieval-based generation of women biographies.arXiv preprint arXiv:2204.05879, 2022

  185. [193]

    Large language models for software engineering: Survey and open problems

    Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M Zhang. Large language models for software engineering: Survey and open problems. In 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (IC...

  186. [194]

    A survey on rag meeting llms: Towards retrieval-augmented large language models

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. A survey on rag meeting llms: Towards retrieval-augmented large language models. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages ...

  187. [195]

    Ai hospital: Benchmarking large language models in a multi-agent medical interaction simulator

    Zhihao Fan, Jialong Tang, Wei Chen, Siyuan Wang, Zhongyu Wei, Jun Xi, Fei Huang, and Jingren Zhou. Ai hospital: Benchmarking large language models in a multi-agent medical interaction simulator. arXiv preprint arXiv:2402.09742, 2024

  188. [196]

    Ai-newton: A concept-driven physical law discovery system without prior physical knowledge.arXiv preprint arXiv:2504.01538, 2025

    You-Le Fang, Dong-Shan Jian, Xiang Li, and Yan-Qing Ma. Ai-newton: A concept-driven physical law discovery system without prior physical knowledge.arXiv preprint arXiv:2504.01538, 2025

  189. [197]

    Shai Farber. Enhancing peer review efficiency: A mixed-methods analysis of artificial intelligence- assisted reviewer selection across academic disciplines.Learned Publishing, 37(4):e1638, Oct 2024

  190. [198]

    Enhancing academic decision-making: A pilot study of ai-supported journal selection in higher education.Innovative Higher Education, pages 1–19, Feb 2025

    Shai Farber. Enhancing academic decision-making: A pilot study of ai-supported journal selection in higher education.Innovative Higher Education, pages 1–19, Feb 2025

  191. [199]

    Wikiatomicedits: A multilingual corpus of wikipedia edits for modeling language and discourse.arXiv preprint arXiv:1808.09422, 2018

    Manaal Faruqui, Ellie Pavlick, Ian Tenney, and Dipanjan Das. Wikiatomicedits: A multilingual corpus of wikipedia edits for modeling language and discourse.arXiv preprint arXiv:1808.09422, 2018

  192. [200]

    Uncovering bottlenecks and optimizing scientific lab workflows with cycle time reduction agents

    Yao Fehlis. Uncovering bottlenecks and optimizing scientific lab workflows with cycle time reduction agents. arXiv preprint arXiv:2505.21534, 2025

  193. [201]

    Accelerating drug discovery with artificial: a whole-lab orchestration and scheduling system for self-driving labs.arXiv preprint arXiv:2504.00986, 2025

    Yao Fehlis, Paul Mandel, Charles Crain, Betty Liu, and David Fuller. Accelerating drug discovery with artificial: a whole-lab orchestration and scheduling system for self-driving labs.arXiv preprint arXiv:2504.00986, 2025

  194. [202]

    A bioactivity foundation model using pairwise meta-learning

    Bin Feng, Zequn Liu, Nanlan Huang, Zhiping Xiao, Haomiao Zhang, Srbuhi Mirzoyan, Hanwen Xu, Jiaran Hao, Yinghui Xu, Ming Zhang, et al. A bioactivity foundation model using pairwise meta-learning. Nature Machine Intelligence, 6(8):962–974, Aug 2024

  195. [203]

    Openfoamgpt 2.0: end-to-end, trustworthy automation for computational fluid dynamics.arXiv preprint arXiv:2504.19338, 2025

    Jingsen Feng, Ran Xu, and Xu Chu. Openfoamgpt 2.0: end-to-end, trustworthy automation for computational fluid dynamics.arXiv preprint arXiv:2504.19338, 2025

  196. [204]

    Sciknoweval: Evaluating multi-level scientific knowledge of large language models.arXiv preprint arXiv:2406.09098, 2024

    Kehua Feng, Keyan Ding, Weijie Wang, Xiang Zhuang, Zeyuan Wang, Ming Qin, Yu Zhao, Jianhua Yao, Qiang Zhang, and Huajun Chen. Sciknoweval: Evaluating multi-level scientific knowledge of large language models.arXiv preprint arXiv:2406.09098, 2024

  197. [205]

    Cocoa: Co-planning and co-execution with ai agents.arXiv preprint arXiv:2412.10999, 2024

    KJ Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S Weld, Amy X Zhang, and Joseph Chee Chang. Cocoa: Co-planning and co-execution with ai agents.arXiv preprint arXiv:2412.10999, 2024

  198. [206]

    Agentic assistant for material scientists

    Ruozhu Feng, Yangang Liang, Tianzhixi Yin, Peiyuan Gao, and Wei Wang. Agentic assistant for material scientists. Apr 2025

  199. [207]

    Grapheval: A lightweight graph-based llm framework for idea evaluation.arXiv preprint arXiv:2503.12600, 2025

    Tao Feng, Yihang Sun, and Jiaxuan You. Grapheval: A lightweight graph-based llm framework for idea evaluation.arXiv preprint arXiv:2503.12600, 2025

  200. [208]

    Surveysum: A dataset for summarizing multiple scientific articles into a survey section

    Leandro Carísio Fernandes, Gustavo Bartz Guedes, Thiago Soares Laitz, Thales Sales Almeida, Rodrigo Nogueira, Roberto Lotufo, and Jayr Pereira. Surveysum: A dataset for summarizing multiple scientific articles into a survey section. InBrazilian Conference on Intelligent System...

  201. [209]

    Fairness and bias in artificial intelligence: A brief survey of sources, impacts, and mitigation strategies.Sci, 6(1):3, Dec 2023

    Emilio Ferrara. Fairness and bias in artificial intelligence: A brief survey of sources, impacts, and mitigation strategies.Sci, 6(1):3, Dec 2023. 64

  202. [210]

    Deepmind and biontech build ai lab assistants for scientific re- search

    Financial Times. Deepmind and biontech build ai lab assistants for scientific re- search. Financial Times , Oct 2024. URL https://www.ft.com/content/ 64b1bb33-095e-4cc5-a911-50df76fa3d1d

  203. [211]

    Baldur: Whole-proof generation and repair with large language models

    Emily First, Markus N Rabe, Talia Ringer, and Yuriy Brun. Baldur: Whole-proof generation and repair with large language models. InProceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pages 1229–124...

  204. [212]

    Marcio Fonseca and Shay Cohen. Can large language model summarizers adapt to diverse scientific communication goals? In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Findings of the Association for Computational Linguistics: ACL 2024, pages 8599–8618, Bangkok, Thailan...

  205. [213]

    Splade v2: Sparse lexical and expansion model for information retrieval.arXiv preprint arXiv:2109.10086, 2021

    Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clinchant. Splade v2: Sparse lexical and expansion model for information retrieval.arXiv preprint arXiv:2109.10086, 2021

  206. [214]

    Explainable and interpretable artificial intelligence in medicine: a systematic bibliometric review.Discover Artificial Intelligence, 4 (1):15, Feb 2024

    Maria Frasca, Davide La Torre, Gabriella Pravettoni, and Ilaria Cutica. Explainable and interpretable artificial intelligence in medicine: a systematic bibliometric review.Discover Artificial Intelligence, 4 (1):15, Feb 2024

  207. [215]

    Data for mathematical copilots: Better ways of presenting proofs for machine learning.arXiv preprint arXiv:2412.15184, 2024

    Simon Frieder, Jonas Bayer, Katherine M Collins, Julius Berner, Jacob Loader, András Juhász, Fabian Ruehle, Sean Welleck, Gabriel Poesia, Ryan-Rhys Griffiths, et al. Data for mathematical copilots: Better ways of presenting proofs for machine learning.arXiv preprint arXiv:2412...

  208. [216]

    Rule-based, neural and llm back-translation: Comparative insights from a variant of ladin.arXiv preprint arXiv:2407.08819, 2024

    Samuel Frontull and Georg Moser. Rule-based, neural and llm back-translation: Comparative insights from a variant of ladin.arXiv preprint arXiv:2407.08819, 2024

  209. [217]

    Peer review expert group recommendation: A multi-subject coverage-based approach.Expert Systems with Applications, 264:125971, Mar 2025

    Yongfan Fu, Jian Luo, Guofang Nan, and Dahui Li. Peer review expert group recommendation: A multi-subject coverage-based approach.Expert Systems with Applications, 264:125971, Mar 2025

  210. [218]

    Intelligent summaries: Will artificial intelligence mark the finale for biomedical literature reviews?Learned Publishing, Dec 2024

    Carlo Galli, Chiara Moretti, and Elena Calciolari. Intelligent summaries: Will artificial intelligence mark the finale for biomedical literature reviews?Learned Publishing, Dec 2024

  211. [219]

    Research- codeagent: An llm multi-agent system for automated codification of research methodologies.arXiv preprint arXiv:2504.20117, 2025

    Shubham Gandhi, Dhruv Shah, Manasi Patwardhan, Lovekesh Vig, and Gautam Shroff. Research- codeagent: An llm multi-agent system for automated codification of research methodologies.arXiv preprint arXiv:2504.20117, 2025

  212. [220]

    Amrita Ganguly, Aditya Johri, Areej Ali, and Nora McDonald. Generative artificial intelligence for academic research: evidence from guidance issued for researchers by higher education institutions in the united states.AI and Ethics, pages 1–17, Mar 2025

  213. [221]

    Grammars of formal uncertainty: When to trust llms in automated reasoning tasks.arXiv preprint arXiv:2505.20047, 2025

    Debargha Ganguly, Vikash Singh, Sreehari Sankar, Biyao Zhang, Xuecen Zhang, Srinivasan Iyengar, Xiaotian Han, Amit Sharma, Shivkumar Kalyanaraman, and Vipin Chaudhary. Grammars of formal uncertainty: When to trust llms in automated reasoning tasks.arXiv preprint arXiv:2505.20047, 2025

  214. [222]

    Cur- rent strategies to address data scarcity in artificial intelligence-based drug discovery: A comprehensive review

    Amit Gangwal, Azim Ansari, Iqrar Ahmad, Abul Kalam Azad, and Wan Mohd Azizi Wan Sulaiman. Cur- rent strategies to address data scarcity in artificial intelligence-based drug discovery: A comprehensive review. Computers in Biology and Medicine, 179:108734, Sep 2024

  215. [223]

    Reviewagents: Bridging the gap between human and ai-generated paper reviews.arXiv preprint arXiv:2503.08506, 2025

    Xian Gao, Jiacheng Ruan, Jingsheng Gao, Ting Liu, and Yuzhuo Fu. Reviewagents: Bridging the gap between human and ai-generated paper reviews.arXiv preprint arXiv:2503.08506, 2025. 65

  216. [224]

    Graph of ai ideas: Leveraging knowledge graphs and llms for ai research idea generation.arXiv preprint arXiv:2503.08549, 2025

    Xian Gao, Zongyun Zhang, Mingye Xie, Ting Liu, and Yuzhuo Fu. Graph of ai ideas: Leveraging knowledge graphs and llms for ai research idea generation.arXiv preprint arXiv:2503.08549, 2025

  217. [225]

    Spaceqa: Answering questions about the design of space missions and space craft concepts

    Andres Garcia-Silva, Cristian Berrio, Jose Manuel Gomez-Perez, Jose Antonio Martínez-Heras, Alessan- dro Donati, and Ilaria Roma. Spaceqa: Answering questions about the design of space missions and space craft concepts. InProceedings of the 45th International ACM SIGIR Confere...

  218. [226]

    Neurodisk: An ai approach to automate continuous inquiry-driven discoveries in neuroimaging genetics.bioRxiv, Feb 2025

    Daniel Garijo, Qifan Yang, Hernán Vargas, Shruti P Gadewar, Kevin Low, Varun Ratnakar, Maximiliano Osorio, Alyssa H Zhu, Agnes McMahon, Yolanda Gil, et al. Neurodisk: An ai approach to automate continuous inquiry-driven discoveries in neuroimaging genetics.bioRxiv, Feb 2025

  219. [227]

    Mir: Methodology inspiration retrieval for scientific research problems.arXiv preprint arXiv:2506.00249, 2025

    Aniketh Garikaparthi, Manasi Patwardhan, Aditya Sanjiv Kanade, Aman Hassan, Lovekesh Vig, and Arman Cohan. Mir: Methodology inspiration retrieval for scientific research problems.arXiv preprint arXiv:2506.00249, 2025

  220. [228]

    Iris: Interactive research ideation system for accelerating scientific discovery.arXiv preprint arXiv:2504.16728, 2025

    Aniketh Garikaparthi, Manasi Patwardhan, Lovekesh Vig, and Arman Cohan. Iris: Interactive research ideation system for accelerating scientific discovery.arXiv preprint arXiv:2504.16728, 2025

  221. [229]

    Baco: A background knowledge-and content-based framework for citing sentence generation

    Yubin Ge, Ly Dinh, Xiaofeng Liu, Jinsong Su, Ziyao Lu, Ante Wang, and Jana Diesner. Baco: A background knowledge-and content-based framework for citing sentence generation. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th I...

  222. [230]

    Human-llm coevolution: Evidence from academic writing.arXiv preprint arXiv:2502.09606, 2025

    Mingmeng Geng and Roberto Trotta. Human-llm coevolution: Evidence from academic writing.arXiv preprint arXiv:2502.09606, 2025

  223. [231]

    How paperpal enhances english writing quality and improves productivity for japanese academics

    Elizabeth Oommen George. How paperpal enhances english writing quality and improves productivity for japanese academics. Aug 2024

  224. [232]

    Sparks: Inspiration for science writing using language models

    Katy Ilonka Gero, Vivian Liu, and Lydia Chilton. Sparks: Inspiration for science writing using language models. InProceedings of the 2022 ACM Designing Interactive Systems Conference, pages 1002–1019, Jun 2022

  225. [233]

    Protagents: protein discovery via large language model multi-agent collaborations combining physics and machine learning.Digital Discovery, 3(7):1389– 1409, May 2024

    Alireza Ghafarollahi and Markus J Buehler. Protagents: protein discovery via large language model multi-agent collaborations combining physics and machine learning.Digital Discovery, 3(7):1389– 1409, May 2024

  226. [234]

    Sciagents: Automating scientific discovery through multi-agent intelligent graph reasoning.arXiv preprint arXiv:2409.05556, 2024

    Alireza Ghafarollahi and Markus J Buehler. Sciagents: Automating scientific discovery through multi-agent intelligent graph reasoning.arXiv preprint arXiv:2409.05556, 2024

  227. [235]

    Sparks: Multi-agent artificial intelligence model discovers protein design principles.arXiv preprint arXiv:2504.19017, 2025

    Alireza Ghafarollahi and Markus J Buehler. Sparks: Multi-agent artificial intelligence model discovers protein design principles.arXiv preprint arXiv:2504.19017, 2025

  228. [236]

    Robin: A multi-agent system for automating scientific discovery.arXiv preprint arXiv:2505.13400

    Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J Szostkiewicz, Jon M Laurent, Muhammed T Razzak, Andrew D White, Michaela M Hinks, and Samuel G Rodriques. Robin: A multi-agent system for automating scientific discovery.arXiv preprint arXiv:2505.13400

  229. [237]

    Hgtdr: Advancing drug repurposing with heterogeneous graph transformers.Bioinformatics, 40(7):btae349, Jul 2024

    Ali Gharizadeh, Karim Abbasi, Amin Ghareyazi, Mohammad RK Mofrad, and Hamid R Rabiee. Hgtdr: Advancing drug repurposing with heterogeneous graph transformers.Bioinformatics, 40(7):btae349, Jul 2024. 66

  230. [238]

    Peer review analyze: A novel benchmark resource for computational analysis of peer reviews.Plos one, 17(1):e0259238, Jan 2022

    Tirthankar Ghosal, Sandeep Kumar, Prabhat Kumar Bharti, and Asif Ekbal. Peer review analyze: A novel benchmark resource for computational analysis of peer reviews.Plos one, 17(1):e0259238, Jan 2022

  231. [239]

    Quaser: Question answering with scalable extractive rationalization

    Asish Ghoshal, Srinivasan Iyer, Bhargavi Paranjape, Kushal Lakhotia, Scott Wen-tau Yih, and Yashar Mehdad. Quaser: Question answering with scalable extractive rationalization. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Informati...

  232. [240]

    Ideasaredimesadozen: Large language models for idea generation in innovation.The Wharton School Research Paper Forthcoming, Jul 2023

    KaranGirotra,LennartMeincke,ChristianTerwiesch,andKarlTUlrich. Ideasaredimesadozen: Large language models for idea generation in innovation.The Wharton School Research Paper Forthcoming, Jul 2023

  233. [241]

    How human–ai feedback loops alter human perceptual, emotional and social judgements.Nature Human Behaviour, 9(2):345–359, 2025

    Moshe Glickman and Tali Sharot. How human–ai feedback loops alter human perceptual, emotional and social judgements.Nature Human Behaviour, 9(2):345–359, 2025

  234. [242]

    Missing counter-evidence renders nlp fact-checking unrealistic for misinformation.arXiv preprint arXiv:2210.13865, 2022

    Max Glockner, Yufang Hou, and Iryna Gurevych. Missing counter-evidence renders nlp fact-checking unrealistic for misinformation.arXiv preprint arXiv:2210.13865, 2022

  235. [243]

    Grounding fallacies misrepresenting scientific publications in evidence.arXiv preprint arXiv:2408.12812, 2024

    Max Glockner, Yufang Hou, Preslav Nakov, and Iryna Gurevych. Grounding fallacies misrepresenting scientific publications in evidence.arXiv preprint arXiv:2408.12812, 2024

  236. [244]

    Hiperrag: High-performance retrieval augmented generation for scientific insights.arXiv preprint arXiv:2505.04846, Jun 2025

    Ozan Gokdemir, Carlo Siebenschuh, Alexander Brace, Azton Wells, Brian Hsu, Kyle Hippe, Priyanka V Setty, Aswathy Ajith, J Gregory Pauloski, Varuni Sastry, et al. Hiperrag: High-performance retrieval augmented generation for scientific insights.arXiv preprint arXiv:2505.04846, Jun 2025

  237. [245]

    Frontiers: Can large language models capture human preferences? Marketing Science, 43(4):709–722, Apr 2024

    Ali Goli and Amandeep Singh. Frontiers: Can large language models capture human preferences? Marketing Science, 43(4):709–722, Apr 2024. doi: 10.1287/mksc.2023.0306. URLhttps://doi. org/10.1287/mksc.2023.0306

  238. [246]

    Catalina Gomez, Sue Min Cho, Shichang Ke, Chien-Ming Huang, and Mathias Unberath. Human-ai collaboration is not very collaborative yet: A taxonomy of interaction patterns in ai-assisted decision making from a systematic review.Frontiers in Computer Science, 6:1521066, Jan 2025

  239. [247]

    Automatic chemical design using a data-driven continuous representation of molecules

    Rafael Gómez-Bombarelli, Jennifer N Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D Hirzel, Ryan P Adams, and Alán Aspuru-Guzik. Automatic chemical design using a data-driven continuous repr...

  240. [248]

    Look, read and enrich-learning from scientific figures and their captions

    Jose Manuel Gomez-Perez and Raul Ortega. Look, read and enrich-learning from scientific figures and their captions. InProceedings of the 10th International Conference on Knowledge Capture, pages 101–108, Sep 2019

  241. [249]

    Rubén González-Sendino, Emilio Serrano, and Javier Bajo. Mitigating bias in artificial intelligence: Fair data generation via causal models for transparent and explainable decision-making.Future Generation Computer Systems, 155:384–401, Jun 2024

  242. [250]

    Lf: a foundational higher-order-logic

    Zachary Goodsell and Juhani Yli-Vakkuri. Lf: a foundational higher-order-logic. arXiv preprint arXiv:2401.11050, 2024

  243. [251]

    Hallucination mitigation using agentic ai natural language-based frameworks

    Diego Gosmar and Deborah A Dahl. Hallucination mitigation using agentic ai natural language-based frameworks. arXiv preprint arXiv:2501.13946, 2025. 67

  244. [252]

    Towards an ai co-scientist.arXiv preprint arXiv:2502.18864, 2025

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al. Towards an ai co-scientist.arXiv preprint arXiv:2502.18864, 2025

  245. [253]

    The llama 3 herd of models

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024

  246. [254]

    Hamiltonian neural networks.Advances in Neural Information Processing Systems, 32, Jul 2019

    Samuel Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian neural networks.Advances in Neural Information Processing Systems, 32, Jul 2019

  247. [255]

    Agentic ai for scientific discovery: A survey of progress, challenges, and future directions.arXiv preprint arXiv:2503.08979, 2025

    Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack. Agentic ai for scientific discovery: A survey of progress, challenges, and future directions.arXiv preprint arXiv:2503.08979, 2025

  248. [256]

    Large language models orchestrating structured reasoning achieve kaggle grand- master level.arXiv preprint arXiv:2411.03562, 2024

    Antoine Grosnit, Alexandre Maraval, James Doran, Giuseppe Paolo, Albert Thomas, Refinath Shahul Hameed Nabeezath Beevi, Jonas Gonzalez, Khyati Khandelwal, Ignacio Iacobacci, Abdelhakim Benechehab, et al. Large language models orchestrating structured reasoning achieve kaggle g...

  249. [257]

    Blade: Benchmarking language model agents for data-driven science

    Ken Gu, Ruoxi Shang, Ruien Jiang, Keying Kuang, Richard-John Lin, Donghe Lyu, Yue Mao, Youran Pan, Teng Wu, Jiaqian Yu, et al. Blade: Benchmarking language model agents for data-driven science. arXiv preprint arXiv:2408.09667, 2024

  250. [258]

    Controllable citation sentence generation with language models

    Nianlong Gu and Richard HR Hahnloser. Controllable citation sentence generation with language models. arXiv preprint arXiv:2211.07066, 2022

  251. [259]

    Llms can realize combinatorial creativity: generating creative ideas via llms for scientific research.arXiv preprint arXiv:2412.14141, 2024

    Tianyang Gu, Jingjin Wang, Zhihao Zhang, and HaoHong Li. Llms can realize combinatorial creativity: generating creative ideas via llms for scientific research.arXiv preprint arXiv:2412.14141, 2024

  252. [260]

    Generation and human-expert evaluation of interesting research ideas using knowledge graphs and large language models.arXiv preprint arXiv:2405.17044, 2024

    Xuemei Gu and Mario Krenn. Generation and human-expert evaluation of interesting research ideas using knowledge graphs and large language models.arXiv preprint arXiv:2405.17044, 2024

  253. [261]

    Ai-assisted drug re-purposing for human liver fibrosis.bioRxiv, pages 2025–04, May 2025

    Yuan Guan, Jakkapong Inchai, Zhuoqing Fang, Jacky Law, Alberto Alonzo Garcia Brito, Annalisa Pawlosky, Juraj GottWeis, Alexander Daryin, Artiom Myaskovsky, Anil Palepu, et al. Ai-assisted drug re-purposing for human liver fibrosis.bioRxiv, pages 2025–04, May 2025

  254. [262]

    A systematic review of federated learning: Challenges, aggregation methods, and development tools.Journal of Network and Computer Applications, 220:103714, Nov 2023

    Badra Souhila Guendouzi, Samir Ouchani, Hiba EL Assaad, and Madeleine EL Zaher. A systematic review of federated learning: Challenges, aggregation methods, and development tools.Journal of Network and Computer Applications, 220:103714, Nov 2023

  255. [263]

    Guided by guardrails: Control barrier functions as safety instructors for robotic learning.arXiv preprint arXiv:2505.18858, 2025

    Maeva Guerrier, Karthik Soma, Hassan Fouad, and Giovanni Beltrame. Guided by guardrails: Control barrier functions as safety instructors for robotic learning.arXiv preprint arXiv:2505.18858, 2025

  256. [264]

    Deepseek-coder: When the large language model meets programming–the rise of code intelligence.arXiv preprint arXiv:2401.14196, 2024

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al. Deepseek-coder: When the large language model meets programming–the rise of code intelligence.arXiv preprint arXiv:2401.14196, 2024

  257. [265]

    Deepseek-r1: Incentivizingreasoningcapabilityinllmsviareinforcement learning

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma,PeiyiWang,XiaoBi,etal. Deepseek-r1: Incentivizingreasoningcapabilityinllmsviareinforcement learning. arXiv preprint arXiv:2501.12948, 2025. 68

  258. [266]

    Ds-agent: Automated data science by empowering large language models with case-based reasoning

    Siyuan Guo, Cheng Deng, Ying Wen, Hechang Chen, Yi Chang, and Jun Wang. Ds-agent: Automated data science by empowering large language models with case-based reasoning. arXiv preprint arXiv:2402.17453, 2024

  259. [267]

    Automated lay language summarization of biomedical scientific reviews

    Yue Guo, Wei Qiu, Yizhong Wang, and Trevor Cohen. Automated lay language summarization of biomedical scientific reviews. InProceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 160–168, May 2021

  260. [268]

    Sciverse: Unveiling the knowledge comprehension and visual reasoning of lmms on multi-modal scientific problems.arXiv preprint arXiv:2503.10627, 2025

    Ziyu Guo, Ray Zhang, Hao Chen, Jialin Gao, Dongzhi Jiang, Jiaze Wang, and Pheng-Ann Heng. Sciverse: Unveiling the knowledge comprehension and visual reasoning of lmms on multi-modal scientific problems.arXiv preprint arXiv:2503.10627, 2025

  261. [269]

    Artificial intelligence to deep learning: machine intelligence approach for drug discovery.Molecular diversity, 25:1315–1360, Apr 2021

    Rohan Gupta, Devesh Srivastava, Mehar Sahu, Swati Tiwari, Rashmi K Ambasta, and Pravir Kumar. Artificial intelligence to deep learning: machine intelligence approach for drug discovery.Molecular diversity, 25:1315–1360, Apr 2021

  262. [270]

    All that glitters is not novel: Plagiarism in ai generated research

    Tarun Gupta and Danish Pruthi. All that glitters is not novel: Plagiarism in ai generated research. arXiv preprint arXiv:2502.16487, 2025

  263. [271]

    Language agents mirror human causal reasoning biases

    Anthony GX-Chen, Dongyan Lin, Mandana Samiei, Doina Precup, Blake A Richards, Rob Fergus, and Kenneth Marino. Language agents mirror human causal reasoning biases. how can we help them think like scientists?arXiv preprint arXiv:2505.09614, 2025

  264. [272]

    Enhance innovation by boosting idea generation with large language models

    Hendrik Haarmann. Enhance innovation by boosting idea generation with large language models. INFORMS Journal on Computing, Jul 2025

  265. [273]

    Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt.Nature Computational Science, 3(10):833–838, October 2023

    Thilo Hagendorff, Sarah Fabi, and Michal Kosinski. Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt.Nature Computational Science, 3(10):833–838, October 2023. ISSN 2662-8457. doi: 10.1038/s43588-023-00527-x. URLhttp...

  266. [274]

    Multi-agent risks from advanced ai.arXiv preprint arXiv:2502.14143, 2025

    Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak, et al. Multi-agent risks from advanced ai.arXiv preprint arXiv:2502.14143, 2025

  267. [275]

    Improving low-resource languages in pre-trained multilingual language models

    Viktor Hangya, Hossain Shaikh Saadi, and Alexander Fraser. Improving low-resource languages in pre-trained multilingual language models. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11993–12006, Jan 2022

  268. [276]

    Ethical and bias considerations in artificial intelligence/machine learning.Modern Pathology, 38(3):100686, Mar 2025

    Matthew G Hanna, Liron Pantanowitz, Brian Jackson, Octavia Palmer, Shyam Visweswaran, Joshua Pantanowitz, Mustafa Deebajah, and Hooman H Rashidi. Ethical and bias considerations in artificial intelligence/machine learning.Modern Pathology, 38(3):100686, Mar 2025

  269. [277]

    Hlm-cite: Hybrid language model workflow for text-based scientific citation prediction.arXiv preprint arXiv:2410.09112, 2024

    Qianyue Hao, Jingyang Fan, Fengli Xu, Jian Yuan, and Yong Li. Hlm-cite: Hybrid language model workflow for text-based scientific citation prediction.arXiv preprint arXiv:2410.09112, 2024

  270. [278]

    Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in Neural Information Processing Systems, 36:45870–45894, Dec 2023

    Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in Neural Information Processing Systems, 36:45870–45894, Dec 2023

  271. [279]

    Airus: a simple workflow for ai-assisted exploration of scientific data.bioRxiv, pages 2025–02, Feb 2025

    Kenneth D Harris. Airus: a simple workflow for ai-assisted exploration of scientific data.bioRxiv, pages 2025–02, Feb 2025. 69

  272. [280]

    Efficacy analysis of online artificial intelligence fact-checking tools.The International Review of Information Ethics, 33(1), Apr 2024

    Russell Hartley. Efficacy analysis of online artificial intelligence fact-checking tools.The International Review of Information Ethics, 33(1), Apr 2024

  273. [281]

    LLM-rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts

    Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, and Chris Kedzie. LLM-rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceedings of the 62nd Annual Meet...

  274. [282]

    Interpreting black-box models: a review on explainable artificial intelligence.Cognitive Computation, 16(1):45–74, Aug 2024

    Vikas Hassija, Vinay Chamola, Atmesh Mahapatra, Abhinandan Singal, Divyansh Goel, Kaizhu Huang, Simone Scardapane, Indro Spinelli, Mufti Mahmud, and Amir Hussain. Interpreting black-box models: a review on explainable artificial intelligence.Cognitive Computation, 16(1):45–74,...

  275. [283]

    Perspective on utilizing foundation models for laboratory automation in materials research.arXiv preprint arXiv:2506.12312, 2025

    Kan Hatakeyama-Sato, Toshihiko Nishida, Kenta Kitamura, Yoshitaka Ushiku, Koichi Takahashi, Yuta Nabae, and Teruaki Hayakawa. Perspective on utilizing foundation models for laboratory automation in materials research.arXiv preprint arXiv:2506.12312, 2025

  276. [284]

    Autonomous robotic system with optical coherence tomography guidance for vascular anastomosis.arXiv preprint arXiv:2410.07493, 2024

    Jesse Haworth, Rishi Biswas, Justin Opfermann, Michael Kam, Yaning Wang, Desire Pantalone, Francis X Creighton, Robin Yang, Jin U Kang, and Axel Krieger. Autonomous robotic system with optical coherence tomography guidance for vascular anastomosis.arXiv preprint arXiv:2410.07493, 2024

  277. [285]

    Simulating 500 million years of evolution with a language model.Science, page eads0018, Jan 2025

    Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q Tran, Jonathan Deaton, Marius Wiggert, et al. Simulating 500 million years of evolution with a language model.Science, page eads0018, Jan 2025

  278. [286]

    From reasoning to learning: A survey on hypothesis discovery and rule learning with large language models.arXiv preprint arXiv:2505.21935, 2025

    Kaiyu He and Zhiyu Chen. From reasoning to learning: A survey on hypothesis discovery and rule learning with large language models.arXiv preprint arXiv:2505.21935, 2025

  279. [287]

    Pasa: An llm agent for comprehensive academic paper search.arXiv preprint arXiv:2501.10120, 2025

    Yichen He, Guanhua Huang, Peiyuan Feng, Yuan Lin, Yuchen Zhang, Hang Li, et al. Pasa: An llm agent for comprehensive academic paper search.arXiv preprint arXiv:2501.10120, 2025

  280. [288]

    Natural language hypotheses in scientific papers and how to tame them: Suggested steps for formalizing complex scientific claims

    Tina Heger, Alsayed Algergawy, Marc Brinner, Jonathan M Jeschke, Birgitta König-Ries, Daniel Mietchen, and Sina Zarrieß. Natural language hypotheses in scientific papers and how to tame them: Suggested steps for formalizing complex scientific claims. InConference on Advances i...

  281. [289]

    Randomized trial of a generative ai chatbot for mental health treatment.Nejm Ai, 2(4):AIoa2400802, Mar 2025

    MichaelVHeinz,DanielMMackin,BriannaMTrudeau,SukanyaBhattacharya,YinzhouWang,HaleyA Banta, Abi D Jewett, Abigail J Salzhauer, Tess Z Griffin, and Nicholas C Jacobson. Randomized trial of a generative ai chatbot for mental health treatment.Nejm Ai, 2(4):AIoa2400802, Mar 2025

  282. [290]

    Literature based discovery: models, methods, and trends.Journal of biomedical informatics, 74:20–32, Oct 2017

    Sam Henry and Bridget T McInnes. Literature based discovery: models, methods, and trends.Journal of biomedical informatics, 74:20–32, Oct 2017

  283. [291]

    Scitrust: Evaluating the trustworthiness of large language models for science

    Emily Herron, Junqi Yin, and Feiyi Wang. Scitrust: Evaluating the trustworthiness of large language models for science. In SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 72–78. IEEE, Nov 2024

  284. [292]

    Evaluatingandtraininglong-contextlargelanguagemodels forquestionansweringonscientificpapers

    LukasHilgert, DanniLiu, andJanNiehues. Evaluatingandtraininglong-contextlargelanguagemodels forquestionansweringonscientificpapers. InSachinKumar, VidhishaBalachandran, ChanYoungPark, 70 Weijia Shi, Shirley Anugrah Hayati, Yulia Tsvetkov, Noah Smith, Hannaneh Hajishirzi, Dongy...

  285. [293]

    Towards automated related work summarization

    Cong Duy Vu Hoang and Min-Yen Kan. Towards automated related work summarization. InColing 2010: Posters, pages 427–435, Aug 2010

  286. [294]

    Malinowski in the age of ai: Can large lan- guagemodelscreateatextgamebasedonananthropologicalclassic? arXivpreprintarXiv:2410.20536 , 2024

    Michael Peter Hoffmann, Jan Fillies, and Adrian Paschke. Malinowski in the age of ai: Can large lan- guagemodelscreateatextgamebasedonananthropologicalclassic? arXivpreprintarXiv:2410.20536 , 2024

  287. [295]

    Aiscivision: A framework for specializinglargemultimodalmodelsinscientificimageclassification

    Brendan Hogan, Anmol Kabra, Felipe Siqueira Pacheco, Laura Greenstreet, Joshua Fan, Aaron Ferber, Marta Ummus, Alecsander Brito, Olivia Graham, Lillian Aoki, et al. Aiscivision: A framework for specializinglargemultimodalmodelsinscientificimageclassification. arXivpreprintarXi...

  288. [296]

    Deconstructinghuman-aicollaboration: Agency,interaction, and adaptation

    SteffenHolterandMennatallahEl-Assady. Deconstructinghuman-aicollaboration: Agency,interaction, and adaptation. InComputer Graphics forum, volume 43, page e15107. Wiley Online Library, Jun 2024

  289. [297]

    Automatic evaluation metrics for artificially generated scientific research.arXiv preprint arXiv:2503.05712, 2025

    Niklas Höpner, Leon Eshuijs, Dimitrios Alivanistos, Giacomo Zamprogno, and Ilaria Tiddi. Automatic evaluation metrics for artificially generated scientific research.arXiv preprint arXiv:2503.05712, 2025

  290. [745]

    URL https://aclanthology.org/2024.acl-long.745/

  291. [2019]

    doi: 10.18653/v1/D19-1371

    Association for Computational Linguistics. doi: 10.18653/v1/D19-1371. URL https:// aclanthology.org/D19-1371/

  292. [2024]

    doi: 10.18653/v1/2024.findings-acl.508

    Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.508. URL https://aclanthology.org/2024.findings-acl.508/

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.