Pith. sign in

REVIEW 4 major objections 5 minor 12 cited by

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A new four-way taxonomy organizes the AI 'Deep Research' boom, and its feature tables rest on vendor-reported numbers that were only partially verified.

desk verdict A useful organizational survey of the Deep Research ecosystem, undermined by contradictory benchmark tables that need reconciliation before the comparative claims can be trusted. read the letter →

arxiv 2506.12594 v1 pith:YUO7CZMC submitted 2025-06-14 cs.AI cs.MA

classification cs.AIcs.MA
keywords DeepResearchlargelanguagemodelsautonomousagentstaxonomyAIsurveytool-useknowledgesynthesisautomation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to establish an organizing framework for the rapidly growing class of AI systems that automate research workflows, which it calls "Deep Research." It analyzes more than 80 commercial and open-source implementations and proposes a hierarchical taxonomy built on four technical dimensions: foundation models and reasoning engines, tool utilization and environmental interaction, task planning and execution control, and knowledge synthesis and output generation. If the framework holds, it gives researchers and developers a common language to compare system architectures, capabilities, and trade-offs across a fragmented ecosystem. The survey also draws conclusions about architectural patterns, benchmark performance, application suitability, and the field's main technical and ethical challenges.

What carries the argument

The key machinery is the four-dimension hierarchical taxonomy itself: foundation models and reasoning engines, tool utilization and environmental interaction, task planning and execution control, and knowledge synthesis and output generation. The taxonomy organizes everything else in the survey, and it is paired with a four-pattern architectural analysis (monolithic, pipeline, multi-agent, hybrid) that explains how systems manage control flow, component coupling, failure propagation, and deployment flexibility.

What would settle it

A controlled evaluation of several systems (e.g., OpenAI/DeepResearch, Gemini/DeepResearch, Perplexity/DeepResearch, and one open-source alternative) on identical tasks from HLE and GAIA, run under the same protocol with repeated trials, would reveal whether the reported margins hold or whether the benchmark-based ranking is an artifact of vendor self-reporting.

Watch

Extended reading notes

Core claim

The central claim is that the diverse landscape of Deep Research systems can be productively organized by a four-dimensional hierarchical taxonomy. Using that taxonomy as a lens, the paper maps the evolution from general-purpose LLM assistants to specialized research systems, identifies four recurring architectural patterns (monolithic, pipeline-based, multi-agent, and hybrid), compares representative systems on benchmarks such as HLE, MMLU, HotpotQA, and GAIA, and evaluates their suitability across academic, enterprise, financial, educational, and personal knowledge-management applications. The intended contribution is both theoretical, a framework for comparing systems, and practical, a roadmap of technical and ethical challenges and future directions.

Load-bearing premise

The survey's comparative tables and performance conclusions rest on vendor-reported benchmark scores and repository documentation that the authors only partially verified with undocumented direct testing; if those reports are inaccurate, the comparisons lose their foundation.

Editorial extensions

If this is right

  • If the taxonomy is adopted, systems can be compared on common dimensions rather than by vendor marketing claims, making capability gaps and trade-offs visible.
  • The architectural pattern analysis gives system builders a decision tool: monolithic for reasoning coherence, pipeline for modularity, multi-agent for parallelization, hybrid for balance.
  • The benchmark comparison establishes a baseline: commercial systems lead on HLE and GAIA, while open-source systems compete on cost, control, and domain-specific optimization.
  • The roadmap of future directions (advanced reasoning, multimodal integration, domain specialization, human-AI collaboration, standardization) identifies what must happen for the field to mature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same four-dimensional taxonomy could extend beyond Deep Research to any AI system that combines a model, tools, and a workflow, giving the broader agent ecosystem a shared vocabulary.
  • The paper's mention of direct testing of public systems, without reporting details, suggests a verification protocol that future surveys could formalize and publish for reproducibility.
  • The four architectural patterns and their failure-propagation characteristics could serve as a practical design heuristic even where benchmark data remain incomplete.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This survey proposes a hierarchical taxonomy of Deep Research systems organized along four technical dimensions (foundation models and reasoning engines, tool utilization and environmental interaction, task planning and execution control, and knowledge synthesis and output generation). It reviews more than 80 commercial and open-source systems, compares them across feature tables and benchmark scores (Tables 1-13), analyzes architectural patterns (monolithic, pipeline, multi-agent, hybrid), discusses implementation technologies, evaluation methodologies, applications, ethical considerations, and future directions, and provides a public repository of resources.

Significance. If its comparative data were reliable, this survey would be a useful reference contribution: it offers broad coverage of a fast-moving field, a clear organizing taxonomy, explicit system inclusion criteria in Section 5.5.1, and honest acknowledgment of absent benchmark evidence in several places (e.g., TREC in Section 5.1.2, FinEval in Section 5.3.2). The taxonomy is an external framing rather than an output of the surveyed systems, so circularity is not a concern. The paper's main weakness is that its analytical core—the cross-system benchmark comparison—is currently undercut by unresolved internal contradictions and by an undocumented verification procedure that is the only stated mechanism for resolving those contradictions.

major comments (4)
  1. [Section 3.3.1, Tables 8-9, Section 5.3.1] The manuscript contains directly contradictory numbers for the same systems and benchmarks. Table 8 reports Grok3Beta MMLU as 79.9%, while Table 9 reports 92.7% for the same system and source [299]; Section 5.3.1 states that OpenAI/DeepResearch averages 72.57% on GAIA, while Table 8 and Table 9 both report 67.36% pass@1 with source [197]. No explanation is given (e.g., different MMLU versions, different GAIA aggregation methods), and no correction is provided. Because these tables and passages constitute the paper's primary comparative evidence, the comparative claim is not currently supportable without reconciliation.
  2. [Section 5.5.3] The stated data-collection method 4, 'Experimental Verification: Where inconsistencies exist, we conducted direct testing of publicly available systems to verify capabilities,' is load-bearing precisely because the manuscript contains the inconsistencies identified above. However, no test dates, system versions, protocols, or results are reported anywhere. A reader therefore cannot determine which of the conflicting numbers is correct, and the paper's own method for resolving such conflicts is unverifiable as written.
  3. [Section 3.3.1, Table 8] The cross-system benchmark comparison is presented as evidence that commercial systems 'generally demonstrate leading performance,' but Table 8 has sparse and non-overlapping entries across systems: no single benchmark is populated for all or even most rows, several rows contain a single score, and the table mixes different benchmark families (HLE, MMLU, HotpotQA, GAIA) without a common evaluation context. The prose conclusion goes beyond what the table can support; claims should be restricted to pairwise or per-benchmark comparisons where data exist.
  4. [Section 5.5.4, Tables 8-12] The manuscript acknowledges in Section 5.5.4 that 'Systems undergo frequent updates, potentially rendering specific benchmark results obsolete,' yet none of the benchmark tables carry evaluation dates or version identifiers. Since the survey's snapshot is dated April 2025 and the ecosystem evolved rapidly during that period, the absence of dates makes it impossible to verify that the reported scores are mutually contemporaneous, which is a necessary condition for the comparisons in Tables 8-12 to be meaningful.
minor comments (5)
  1. [Section 3.3.1] There is a typo: 'HLE [212] hich measures' should be 'which measures'.
  2. [Section 1.4] The roadmap in Section 1.4 does not match the actual section numbering: implementation technologies are presented in Section 4, evaluation methodologies in Section 5, applications in Section 6, ethical considerations in Section 7, and future directions in Section 8, whereas the text refers to 'implementation technologies (Section 5), evaluation methodologies (Section 6), applications and use cases (Section 7), ethical considerations (Section 8), and future directions (Section 9)'.
  3. [Table 8] The footnote for GAIA ('GAIA Score(pass@1): Average score') is unclear; GAIA is typically reported as pass@1 averaged over the three difficulty levels, and the label should be aligned with the reporting convention used in the GAIA paper and in Table 9.
  4. [Table 12, Section 5.2.1] Response-time figures are inconsistent: Table 12 gives OpenAI/DeepResearch a range of 5-30 minutes while Section 5.2.1 says '5-10 minutes,' and Table 12 gives Perplexity/DeepResearch 2m59s while Section 5.2.1 says '2-5 minutes'; these should be reconciled or qualified by task complexity.
  5. [Table 8] Table 8 includes rows for Gemini-2.5 and Gemini-2.0-Flash, but Section 1.1's inclusion criteria target Deep Research systems, and the table does not clarify whether these rows refer to Gemini/DeepResearch or to the underlying models; this should be stated explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's taxonomy is an external organizing framework, not a quantity derived from or fitted to the systems it surveys.

full rationale

This paper is a survey, not a predictive derivation: it proposes a four-dimensional taxonomy for describing Deep Research systems and then applies that taxonomy to organize architectural patterns, benchmarks, applications, and challenges. There is no equation, fitted parameter, or benchmark-derived quantity that is later relabeled as a prediction. The taxonomy is defined in Section 2 from technical capabilities, and the surveyed systems are then described in those terms; this is classification rather than derivation, so the central claim does not reduce to its inputs by construction. The paper does not rely on a load-bearing self-citation chain: no 'uniqueness theorem' or prior author result is invoked to make a choice forced, and the GitHub resource link is supporting material, not the foundation of the argument. The internal benchmark discrepancies noted by the skeptical reader (e.g., Grok3Beta MMLU 79.9% vs. 92.7%; GAIA 67.36% vs. 72.57%) are data-reliability or correctness concerns about vendor-reported metrics and undeclared direct testing, not circularity under any of the enumerated patterns. The paper itself flags data limitations and states that inconsistencies were checked, but even a failure to document those checks would affect empirical reliability, not circularity. Accordingly, an honest non-finding is appropriate: the survey is self-contained as an organizing framework, and no circular step can be exhibited from its text.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The survey rests on hand-chosen selection criteria, vendor-reported data, and an asserted taxonomy. No new entities, forces, or fitted physical parameters are introduced.

free parameters (1)
  • System inclusion criteria = At least 2 of 3 core dimensions; public documentation; active development within past 12 months; representational…
    Hand-chosen thresholds in Section 5.5.1 shape the survey scope; changing them alters which systems are analyzed and the resulting comparison.
assumptions (3)
  • domain assumption Vendor-reported benchmarks and documentation accurately represent system capabilities
    Sections 3.3 and 5.5.3 use published scores and repository docs as the primary evidence; the authors did not independently reproduce most claims.
  • ad hoc to paper The four proposed technical dimensions are the fundamental axes for categorizing Deep Research systems
    Section 1.4 and Section 2 assert the taxonomy without deriving it from data or validating inter-rater agreement.
  • domain assumption Benchmark scores from different evaluations are comparable across systems despite differing protocols
    Tables 8 and 9 compare HLE, MMLU, GAIA and other scores without standardizing evaluation conditions; Section 5.5.4 acknowledges this limit.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications." pith.science (2026). https://pith.science/paper/YUO7CZMC

@misc{pith2026250612594,
  author       = {Pith},
  title        = {Pith review of: A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YUO7CZMC}},
  note         = {Machine review of arXiv:2506.12594}
}
read the original abstract

This survey examines the rapidly evolving field of Deep Research systems -- AI-powered applications that automate complex research workflows through the integration of large language models, advanced information retrieval, and autonomous reasoning capabilities. We analyze more than 80 commercial and non-commercial implementations that have emerged since 2023, including OpenAI/Deep Research, Gemini/Deep Research, Perplexity/Deep Research, and numerous open-source alternatives. Through comprehensive examination, we propose a novel hierarchical taxonomy that categorizes systems according to four fundamental technical dimensions: foundation models and reasoning engines, tool utilization and environmental interaction, task planning and execution control, and knowledge synthesis and output generation. We explore the architectural patterns, implementation approaches, and domain-specific adaptations that characterize these systems across academic, scientific, business, and educational applications. Our analysis reveals both the significant capabilities of current implementations and the technical and ethical challenges they present regarding information accuracy, privacy, intellectual property, and accessibility. The survey concludes by identifying promising research directions in advanced reasoning architectures, multimodal integration, domain specialization, human-AI collaboration, and ecosystem standardization that will likely shape the future evolution of this transformative technology. By providing a comprehensive framework for understanding Deep Research systems, this survey contributes to both the theoretical understanding of AI-augmented knowledge work and the practical development of more capable, responsible, and accessible research technologies. The paper resources can be viewed at https://github.com/scienceaix/deepresearch.

Figures

Figures reproduced from arXiv: 2506.12594 by the authors.

Figure 1
Figure 1. Evolution Timeline of Deep Research Systems [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Hierarchical Technical Framework of Deep Research Systems [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Implementation Architecture of Deep Research Systems [PITH_FULL_IMAGE:figures/full_fig_p022_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Monolithic Deep Research Architecture ∙ Sequential Component Organization: Research tasks flow through a predefined sequence of specialized processing modules ∙ Standardized Interfaces: Clear data transformation specifications between pipeline stages enable modular com…
Figure 5
Figure 5. Figure 5: Pipeline-Based Deep Research Architecture [PITH_FULL_IMAGE:figures/full_fig_p024_5.png]
Figure 6
Figure 6. Figure 6: Multi-Agent Deep Research Architecture 4.1.4 Hybrid Architecture Pattern. Hybrid architectures combine elements from multiple architectural patterns to balance their respective advantages within unified implementations. As shown in [PITH_FULL_IMAGE:figures/full_fig_p0…
Figure 7
Figure 7. Figure 7: Hybrid Deep Research Architecture gathering and specialized processing pipelines to achieve sophisticated research capabilities with balanced performance characteristics. 4.1.5 Emerging Agent Framework Ecosystems. Beyond the core architectural patterns described above,…
Figure 8
Figure 8. Figure 8: Multi-dimensional Evaluation Framework for Deep Research Systems [PITH_FULL_IMAGE:figures/full_fig_p035_8.png]
Figure 9
Figure 9. Figure 9: Deep Research Application Domains and Use Cases [PITH_FULL_IMAGE:figures/full_fig_p044_9.png]
Figure 10
Figure 10. Figure 10: Ethical Dimensions of Deep Research Systems [PITH_FULL_IMAGE:figures/full_fig_p057_10.png]
Figure 11
Figure 11. Figure 11: Research Directions for Deep Research Systems [PITH_FULL_IMAGE:figures/full_fig_p065_11.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

    cs.CL 2025-09 conditional novelty 7.0 of 10

    A 7B LLM agent trained with student-led distillation and one-step teacher corrections nearly matches a 72B teacher on reasoning and tool-use benchmarks.

  2. HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research

    cs.IR 2026-07 conditional novelty 6.5 of 10

    A hierarchical evidence-graph benchmark reveals that multimodal deep-research models write fluent reports while failing citation, claim, and answer grounding.

  3. From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

    cs.AI 2026-08 conditional novelty 6.0 of 10

    An automatic pipeline evolves simple QA questions into 500 deep-research tasks with DAG-structured, fact-grounded rubrics that discriminate between models.

  4. Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis

    cs.CL 2026-07 conditional novelty 6.0 of 10

    LLM deep-research agents rarely use historical analogies; a structural-decomposition plus cross-analogy-confirmation agent (CANA) sharply increases mechanism-grounded analogy claims and hidden-factor hits on the new A...

  5. DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation

    cs.CL 2025-12 conditional novelty 6.0 of 10

    DEER uses 7 evaluation dimensions, 101 rubric items, task-specific expert guidance, and unsupported-claim backtracking to score deep-research reports; current systems score lowest on fulfilling expert requests and ana...

  6. SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

    cs.AI 2025-09 conditional novelty 6.0 of 10

    An automated red-teaming framework and 300-case benchmark show that a single unreliable website can induce unsafe responses in LLM search agents, with attack success rates up to 90.5%.

  7. Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG

    cs.CL 2025-09 conditional novelty 6.0 of 10

    In multilingual retrieval-augmented generation, models cite English evidence more accurately than translated evidence, and this language preference can outweigh document relevance.

  8. SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A 20B autonomously reasoning deep-research agent trained with synthetic-data RL reaches 28.7% on Humanity's Last Exam, exceeding several larger and proprietary baselines.

  9. Characterizing Deep Research: A Benchmark and Formal Definition

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Deep research is characterized by high search and reasoning intensity; the new LiveDRBench measures claim-level precision and recall, where the best current model scores 0.55 F1.

  10. FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

    cs.CL 2026-07 conditional novelty 5.5 of 10

    A multi-LLM consensus pipeline turns 14,450 auto-generated candidate rubrics into 2,600 distinguishable gold rubrics that rank 10 financial deep-research systems from 58.58% to 22.23% pass rate.

  11. DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent

    cs.AI 2026-03 conditional novelty 5.0 of 10

    A synthetic benchmark of 9,000 multi-hop web-research questions with difficulty tiers and teacher-generated search trajectories, plus an open-source RL training framework that reportedly lets 3B-parameter agents beat ...

  12. Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory

    cs.AI 2025-10 conditional novelty 5.0 of 10

    Branch-and-Browse, a tree-structured web agent with page-level action memory, reports 35.8% success and up to 40.4% less time than the cited Tree Search baseline on WebArena.

Reference graph

Works this paper leans on

298 extracted references · 38 canonical work pages · cited by 12 Pith papers

  1. [299]

    Sarah Welsh. 2025. AI Benchmark Deep Dive: Gemini 2.5 and Humanity’s Last Exam. https://arize.com/blog/ai- benchmark-deep-dive-gemini-humanitys-last-exam/

  2. [197]

    OpenAI. 2025. Introducing Deep Research. https://openai.com/index/introducing-deep-research/

  3. [1]

    Adilzhan Adilkhanov, Amir Yelenov, Assylkhan Seitzhanov, Ayan Mazhitov, Azamat Abdikarimov, Danissa Sandyk- bayeva, Daryn Kenzhebek, Dinmukhammed Mukashev, Ilyas Umurbekov, Jabrail Chumakov, Kamila Spanova, Karina Burunchina, Madina Yergibay, Margulan Issa, Moldir Zabirova, Nurdaulet Zhuzbay, Nurlan Kabdyshev, Nurlan Zhani- yar, Rasul Yermagambet, Rustam ...

  4. [2]

    Agent-RL. 2024. ReSearch. https://github.com/Agent-RL/ReSearch

  5. [3]

    Agno-AGI. 2025. Agno. https://github.com/agno-agi/agno

  6. [4]

    Garima Agrawal, Sashank Gummuluri, and Cosimo Spera. 2024. Beyond-RAG: Question Identification and Answer Generation in Real-Time Conversations. arXiv:2410.10136 [cs.CL] https://arxiv.org/abs/2410.10136

  7. [5]

    Flowise AI. 2023. Flowise: Low-code LLM Application Building Tool. https://flowiseai.com/

  8. [6]

    Nawaf Alampara, Mara Schilling-Wilhelmi, Martiño Ríos-García, Indrajeet Mandal, Pranav Khetarpal, Hargun Singh Grover, N. M. Anoop Krishnan, and Kevin Maik Jablonka. 2025. Probing the limitations of multimodal language models for chemistry and materials research. arXiv:2411.16955 [cs.LG] https://arxiv.org/abs/2411.16955

Show all 298 references
  1. [7]

    AlphaProof and AlphaGeometry teams. 2024. AI achieves silver-medal standard solving International Mathematical Olympiad problems. https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/

  2. [8]

    Salaheddin Alzubi, Creston Brooks, Purva Chiniya, Edoardo Contente, Chiara von Gerlach, Lucas Irwin, Yihan Jiang, Arda Kaz, Windsor Nguyen, Sewoong Oh, Himanshu Tyagi, and Pramod Viswanath. 2025. Open Deep Search: Democratizing Search with Open-source Reasoning Agents. arXiv:2...

  3. [9]

    Lucio Anderlini, Matteo Barbetti, Giulio Bianchini, Diego Ciangottini, Stefano Dal Pra, Diego Michelotto, Carmelo Pellegrino, Rosa Petrini, Alessandro Pascolini, and Daniele Spiga. 2025. Supporting the development of Machine Learning for fundamental science in a federated Clou...

  4. [10]

    Mehrad Ansari and Seyed Mohamad Moosavi. 2023. Agent-based Learning of Materials Datasets from Scientific Literature. https://github.com/AI4ChemS/Eunomia. arXiv:2312.11690 [cs.AI] https://arxiv.org/abs/2312.11690

  5. [11]

    Anthropic. 2024. Building effective agents. https://www.anthropic.com/engineering/building-effective-agents

  6. [12]

    Antropic. 2024. Model Context Protocol (MCP). https://docs.anthropic.com/en/docs/agents-and-tools/mcp

  7. [13]

    Antropic. 2025. Claude takes research to new places. https://www.anthropic.com/news/research

  8. [14]

    Prakash Aryan. 2024. LLMs as Debate Partners: Utilizing Genetic Algorithms and Adversarial Search for Adaptive Arguments. arXiv:2412.06229 [cs.AI] https://arxiv.org/abs/2412.06229

  9. [15]

    Johnson, Casey Dugan, and Michelle Bachman

    Zahra Ashktorab, Qian Pan, Werner Geyer, Michael Desmond, Marina Danilevsky, James M. Johnson, Casey Dugan, and Michelle Bachman. 2024. Emerging Reliance Behaviors in Human-AI Text Generation: Hallucinations, Data Quality Assessment, and Cognitive Forcing Functions. arXiv:2409...

  10. [16]

    assafelovic. 2023. GPT-Researcher. https://github.com/assafelovic/gpt-researcher/

  11. [17]

    Ahmet Yasin Aytar, Kemal Kilic, and Kamer Kaya. 2024. A Retrieval-Augmented Generation Framework for Academic Literature Navigation in Data Science. arXiv:2412.15404 [cs.IR] https://arxiv.org/abs/2412.15404

  12. [18]

    Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. 2025. ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. arXiv:2404.07738 [cs.CL] https://arxiv.org/ abs/2404.07738

  13. [19]

    Dzmitry Bahdanau, Nicolas Gontier, Gabriel Huang, Ehsan Kamalloo, Rafael Pardinas, Alex Piché, Torsten Scholak, Oleh Shliazhko, Jordan Prince Tremblay, Karam Ghanem, Soham Parikh, Mitul Tiwari, and Quaizar Vohra. 2024. TapeAgents: a Holistic Framework for Agent Development and...

  14. [20]

    Gal Bakal, Ali Dasdan, Yaniv Katz, Michael Kaufman, and Guy Levin. 2025. Experience with GitHub Copilot for Developer Productivity at Zoominfo. arXiv:2501.13282 [cs.SE] https://arxiv.org/abs/2501.13282

  15. [21]

    Howard Balshem, Mark Helfand, Holger J Schünemann, Andrew D Oxman, Regina Kunzand Jan Brozek, Gunn E Vist, Yngve Falck-Ytter, Joerg Meerpohl, Susan Norris, and Gordon H Guyatt. 2011. GRADE guidelines: 3. Rating the quality of evidence. https://pubmed.ncbi.nlm.nih.gov/21208779/

  16. [22]

    Samuel Barham, Orion Weller, Michelle Yuan, Kenton Murray, Mahsa Yarmohammadi, Zhengping Jiang, Siddharth Vashishtha, Alexander Martin, Anqi Liu, Aaron Steven White, Jordan Boyd-Graber, and Benjamin Van Durme

  17. [23]

    Rhea Basappa, Mustafa Tekman, Hong Lu, Benjamin Faught, Sandeep Kakar, and Ashok K. Goel. 2024.Social AI Agents Too Need to Explain Themselves. Springer Nature Switzerland, 351–360. doi:10.1007/978-3-031-63028-6_29

  18. [24]

    Joeran Beel, Min-Yen Kan, and Moritz Baumgart. 2025. Evaluating Sakana’s AI Scientist for Autonomous Research: Wishful Thinking or an Emerging Reality Towards ’Artificial Research Intelligence’ (ARI)? arXiv:2502.14297 [cs.IR] https://arxiv.org/abs/2502.14297

  19. [25]

    Morad Behandish, John Maxwell III, and Johan de Kleer. 2022. AI Research Associate for Early-Stage Scientific Discovery. arXiv:2202.03199 [cs.AI] https://arxiv.org/abs/2202.03199

  20. [26]

    Yoshua Bengio, Michael Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, Sören Mindermann, Adam Oberman, Jesse Richardson, Oliver Richardson, Marc-Antoine Rondeau, Pierre-Luc St-Charles, and David Williams-King. 2025. Superintelligent Agents Pose Catas...

  21. [27]

    Karim Benharrak, Tim Zindulka, and Daniel Buschek. 2024. Deceptive Patterns of Intelligent and Interactive Writing Assistants. arXiv:2404.09375 [cs.HC] https://arxiv.org/abs/2404.09375

  22. [28]

    Zhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, and Huajun Chen. 2024. OceanGPT: A Large Language Model for Ocean Science Tasks. http://oceangpt.zjukg.cn/. arXiv:2310.02031 [cs.CL] https: //arxiv.org/abs/2310.02031

  23. [29]

    Stefano Bianchini, Moritz Müller, and Pierre Pelletier. 2024. Drivers and Barriers of AI Adoption and Use in Scientific Research. arXiv:2312.09843 [cs.CY] https://arxiv.org/abs/2312.09843

  24. [30]

    bindAI. 2025. ChatGPT Deep Research vs Perplexity – Which One Is Better? https://blog.getbind.co/2025/02/03/ chatgpt-deep-research-is-it-better-than-perplexity/

  25. [31]

    Francisco Bolanos, Angelo Salatino, Francesco Osborne, and Enrico Motta. 2024. Artificial Intelligence for Literature Reviews: Opportunities and Challenges. arXiv:2402.08565 [cs.AI] https://arxiv.org/abs/2402.08565

  26. [32]

    Bolt. 2024. Bolt. https://bolt.new/

  27. [33]

    bracai. 2025. MMLU benchmark: Testing LLMs multi-task capabilities. https://www.bracai.eu/post/mmlu-benchmark

  28. [34]

    Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. 2023. ChemCrow: Augmenting large-language models with chemistry tools. arXiv:2304.05376 [physics.chem-ph] https://arxiv.org/abs/ 2304.05376

  29. [35]

    Chris Brown and Jason Cusati. 2024. Exploring the Evidence-Based Beliefs and Behaviors of LLM-Based Programming Assistants. arXiv:2407.13900 [cs.SE] https://arxiv.org/abs/2407.13900

  30. [36]

    browserbase. 2025. Open-operator. https://github.com/browserbase/open-operator

  31. [37]

    btahir. 2024. open_deep_research. https://github.com/btahir/open-deep-research

  32. [38]

    ByteDance. 2024. Coze Space. https://www.coze.cn/space-preview

  33. [39]

    ByteDance. 2025. agent-tars. https://github.com/bytedance/UI-TARS-desktop/tree/main/apps/agent-tars

  34. [40]

    Beatriz Cabrero-Daniel, Tomas Herda, Victoria Pichler, and Martin Eder. 2024. Exploring Human-AI Collaboration in Agile: Customised LLM Meeting Assistants. arXiv:2404.14871 [cs.SE] https://arxiv.org/abs/2404.14871

  35. [41]

    Filipe Calegario, Vanilson Burégio, Francisco Erivaldo, Daniel Moraes Costa Andrade, Kailane Felix, Nathalia Barbosa, Pedro Lucas da Silva Lucena, and César França. 2023. Exploring the intersection of Generative AI and Software Development. arXiv:2312.14262 [cs.SE] https://arx...

  36. [42]

    Nicholas Camara. 2025. open-deep-research. https://github.com/nickscamara/open-deep-research

  37. [43]

    Camel AI. 2025. OWL. https://github.com/camel-ai/owl

  38. [44]

    Mustafa Rafique, Eliu Huerta, Bo Li, Ian Foster, and Rick Stevens

    Franck Cappello, Sandeep Madireddy, Robert Underwood, Neil Getty, Nicholas Lee-Ping Chia, Nesar Ramachandra, Josh Nguyen, Murat Keceli, Tanwi Mallick, Zilinghan Li, Marieme Ngom, Chenhui Zhang, Angel Yanguas-Gil, Evan Antoniuk, Bhavya Kailkhura, Minyang Tian, Yufeng Du, Yuan-S...

  39. [45]

    Peter Cardon, Carolin Fleischmann, Jolanta Aritz, Minna Logemann, and Jeanette Heidewald. 2023. The Challenges and Opportunities of AI-Assisted Writing: Developing AI Literacy for the AI Age. https://journals.sagepub.com/doi/ abs/10.1177/23294906231176517

  40. [46]

    Pan, Shuyi Yang, Lakshya A

    Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica. 2025. Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657 [cs.AI] h...

  41. [47]

    Eric Chamoun, Michael Schlichktrull, and Andreas Vlachos. 2024. Automated Focused Feedback Generation for Scientific Writing Assistance. arXiv:2405.20477 [cs.CL] https://arxiv.org/abs/2405.20477

  42. [48]

    Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate. https://github.com/thunlp/ChatEval. arXiv:2308.07201 [cs.CL] https://arxiv.org/abs/2308.07201 A...

  43. [49]

    Jiangjie Chen, Xintao Wang, Rui Xu, Siyu Yuan, Yikai Zhang, Wei Shi, Jian Xie, Shuang Li, Ruihan Yang, Tinghui Zhu, Aili Chen, Nianqi Li, Lida Chen, Caiyu Hu, Siye Wu, Scott Ren, Ziquan Fu, and Yanghua Xiao. 2024. From Persona to Personalization: A Survey on Role-Playing Langu...

  44. [50]

    Kexin Chen, Hanqun Cao, Junyou Li, Yuyang Du, Menghao Guo, Xin Zeng, Lanqing Li, Jiezhong Qiu, Pheng Ann Heng, and Guangyong Chen. 2024. An Autonomous Large Language Model Agent for Chemical Literature Data Mining. arXiv:2402.12993 [cs.IR] https://arxiv.org/abs/2402.12993

  45. [51]

    Pengcheng Chen, Jin Ye, Guoan Wang, Yanjun Li, Zhongying Deng, Wei Li, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, Benyou Wang, Shaoting Zhang, Bin Fu, Jianfei Cai, Bohan Zhuang, Eric J Seibel, Junjun He, and Yu Qiao

  46. [52]

    Tingting Chen, Srinivas Anumasa, Beibei Lin, Vedant Shah, Anirudh Goyal, and Dianbo Liu. 2025. Auto- Bench: An Automated Benchmark for Scientific Discovery in LLMs. https://github.com/AutoBench/AutoBench. arXiv:2502.15224 [cs.LG] https://arxiv.org/abs/2502.15224

  47. [53]

    Valerie Chen, Alan Zhu, Sebastian Zhao, Hussein Mozannar, David Sontag, and Ameet Talwalkar. 2025. Need Help? Designing Proactive AI Assistants for Programming. arXiv:2410.04596 [cs.HC] https://arxiv.org/abs/2410.04596

  48. [54]

    Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. 2023. AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent B...

  49. [55]

    Qinyuan Cheng, Tianxiang Sun, Xiangyang Liu, Wenwei Zhang, Zhangyue Yin, Shimin Li, Linyang Li, Zhengfu He, Kai Chen, and Xipeng Qiu. 2024. Can AI Assistants Know What They Don’t Know? arXiv:2401.13275 [cs.CL] https://arxiv.org/abs/2401.13275

  50. [56]

    ZhaoCheng, DianeWan, MatthewAbueg,Sahra Ghalebikesabi, RenYi, EugeneBagdasarian,Borja Balle,Stefan Mellem, and Shawn O’Banion. 2024. CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data. https: //www.aimodels.fyi/papers/arxiv/ci-bench-benchmarking-con...

  51. [57]

    Hen- ley

    Bhavya Chopra, Ananya Singha, Anna Fariha, Sumit Gulwani, Chris Parnin, Ashish Tiwari, and Austin Z. Hen- ley. 2023. Conversational Challenges in AI-Powered Data Science: Obstacles, Needs, and Design Opportunities. arXiv:2310.16164 [cs.HC] https://arxiv.org/abs/2310.16164

  52. [58]

    Daniel J. H. Chung, Zhiqi Gao, Yurii Kvasiuk, Tianyi Li, Moritz Münchmeyer, Maja Rudolph, Frederic Sala, and Sai Chaitanya Tadepalli. 2025. Theoretical Physics Benchmark (TPBench) – a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics. https://tpbench.org/. ...

  53. [59]

    Umut Cihan, Vahid Haratian, Arda İçöz, Mert Kaan Gül, Ömercan Devran, Emircan Furkan Bayendur, Baykal Mehmet Uçar, and Eray Tüzün. 2024. Automated Code Review In Practice. arXiv:2412.18531 [cs.SE] https://arxiv.org/abs/ 2412.18531

  54. [60]

    Dave Citron. 2025. Deep Research is now available on Gemini 2.5 Pro Experimental. https://blog.google/products/ gemini/deep-research-gemini-2-5-pro-experimental/

  55. [61]

    Cline. 2024. Cline. https://github.com/cline/cline

  56. [62]

    2025.Devin.ai

    Cognition Labs. 2025.Devin.ai. https://devin.ai

  57. [63]

    Consensus. 2025. Consensus. https://consensus.app/

  58. [64]

    crewAIInc. 2023. CrewAI. https://github.com/crewAIInc/crewAI

  59. [65]

    Cursor. 2023. Cursor. https://www.cursor.com/

  60. [66]

    Danesh, Tu Trinh, Benjamin Plaut, and Nguyen X

    Mohamad H. Danesh, Tu Trinh, Benjamin Plaut, and Nguyen X. Khanh. 2025. Learning to Coordinate with Experts. https://github.com/modanesh/YRC-Bench. arXiv:2502.09583 [cs.LG] https://arxiv.org/abs/2502.09583

  61. [67]

    Kristin M. de Payrebrune, Kathrin Flaßkamp, Tom Ströhla, Thomas Sattel, Dieter Bestle, Benedict Röder, Peter Eberhard, Sebastian Peitz, Marcus Stoffel, Gulakala Rutwik, Borse Aditya, Meike Wohlleben, Walter Sextro, Maximilian Raff, C. David Remy, Manish Yadav, Merten Stender, ...

  62. [68]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, et al. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12...

  63. [69]

    Akash Dhruv and Anshu Dubey. 2025. Leveraging Large Language Models for Code Translation and Software Development in Scientific Computing. arXiv:2410.24119 [cs.SE] https://arxiv.org/abs/2410.24119 84 Xu et al

  64. [70]

    Talissa Dreossi. 2025. Bridging Logic Programming and Deep Learning for Explainability through ILASP.Electronic Proceedings in Theoretical Computer Science416 (Feb. 2025), 314–323. doi:10.4204/eptcs.416.31

  65. [71]

    It makes you think

    Ian Drosos, Advait Sarkar, Xiaotong Xu, and Neil Toronto. 2025. "It makes you think": Provocations Help Restore Critical Thinking to AI-Assisted Knowledge Work. arXiv:2501.17247 [cs.HC] https://arxiv.org/abs/2501.17247

  66. [72]

    Omer Dunay, Daniel Cheng, Adam Tait, Parth Thakkar, Peter C Rigby, Andy Chiu, Imad Ahmad, Arun Ganesan, Chandra Maddila, Vijayaraghavan Murali, Ali Tayyebi, and Nachiappan Nagappan. 2024. Multi-line AI-assisted Code Authoring. arXiv:2402.04141 [cs.SE] https://arxiv.org/abs/2402.04141

  67. [73]

    Steffen Eger, Yong Cao, Jennifer D’Souza, Andreas Geiger, Christian Greisinger, Stephanie Gross, Yufang Hou, Brigitte Krenn, Anne Lauscher, Yizhi Li, Chenghua Lin, Nafise Sadat Moosavi, Wei Zhao, and Tristan Miller. 2025. Transforming Science with Large Language Models: A Surv...

  68. [74]

    Elicit. 2025. Elicit. https://elicit.com/?redirected=true

  69. [75]

    Michael D. Ernst. 2017. Natural Language is a Programming Language: Applying Natural Language Processing to Software Development. https://drops.dagstuhl.de/storage/00lipics/lipics-vol071-snapl2017/LIPIcs.SNAPL.2017.4/ LIPIcs.SNAPL.2017.4.pdf

  70. [76]

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models. arXiv:2405.06211 [cs.CL] https://arxiv.org/abs/2405.06211

  71. [77]

    Flowith. 2024. Flowith Oracle Mode. https://flowith.net/

  72. [78]

    Forethought-Technologies. 2023. AutoChain. https://github.com/Forethought-Technologies/AutoChain

  73. [79]

    César França. 2023. AI empowering research: 10 ways how science can benefit from AI. arXiv:2307.10265 [cs.GL] https://arxiv.org/abs/2307.10265

  74. [80]

    Future-House. 2023. PaperQA. https://github.com/Future-House/paper-qa

  75. [81]

    GAIR-NLP. 2025. DeepResearcher. https://github.com/GAIR-NLP/DeepResearcher

  76. [82]

    Difei Gao, Lei Ji, Luowei Zhou, Kevin Qinghong Lin, Joya Chen, Zihan Fan, and Mike Zheng Shou. 2023. AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn. arXiv:2306.08640 [cs.CV] https: //arxiv.org/abs/2306.08640

  77. [83]

    Shanghua Gao, Ada Fang, Yepeng Huang, Valentina Giunchiglia, Ayush Noori, Jonathan Richard Schwarz, Yasha Ektefaie, Jovana Kondic, and Marinka Zitnik. 2024. Empowering Biomedical Discovery with AI Agents. arXiv:2404.02831 [cs.AI] https://arxiv.org/abs/2404.02831

  78. [84]

    Alireza Ghafarollahi and Markus J. Buehler. 2024. SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning. arXiv:2409.05556 [cs.AI] https://arxiv.org/abs/2409.05556

  79. [85]

    Luca Gioacchini, Marco Mellia, Idilio Drago, Alexander Delsanto, Giuseppe Siracusano, and Roberto Bifulco. 2024. AutoPenBench: Benchmarking Generative Agents for Penetration Testing. https://github.com/lucagioacchini/auto- pen-bench. arXiv:2410.03225 [cs.CR] https://arxiv.org/...

  80. [86]

    Github. 2021. Github Copilot. https://github.com/features/copilot?ref=nav.poetries.top

  81. [87]

    Amr Gomaa, Michael Sargious, and Antonio Krüger. 2024. AdaptoML-UX: An Adaptive User-centered GUI-based AutoML Toolkit for Non-AI Experts and HCI Researchers. https://github.com/MichaelSargious/AdaptoML_UX. arXiv:2410.17469 [cs.HC] https://arxiv.org/abs/2410.17469

  82. [88]

    Google. 2021. BIG-bench. https://github.com/google/BIG-bench

  83. [89]

    Google. 2024. Try Deep Research and our new experimental model in Gemini, your AI assistant. https://blog.google/ products/gemini/google-gemini-deep-research/

  84. [90]

    Google. 2025. A2A. https://github.com/google/A2A

  85. [91]

    Google. 2025. Agent Development Kit. https://google.github.io/adk-docs/

  86. [92]

    Google. 2025. Announcing the Agent2Agent Protocol (A2A). https://developers.googleblog.com/en/a2a-a-new-era-of- agent-interoperability/

  87. [93]

    Google. 2025. Gemini 2.0 Flash (Feb ’25): Intelligence, Performance and Price Analysis. https://artificialanalysis.ai/ models/gemini-2-0-flash

  88. [94]

    Google. 2025. gemini-fullstack-langgraph-quickstart. https://github.com/google-gemini/gemini-fullstack-langgraph- quickstart

  89. [95]

    Google. 2025. NotebookLm. https://notebooklm.google/

  90. [96]

    Kanika Goswami, Puneet Mathur, Ryan Rossi, and Franck Dernoncourt. 2025. ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution. arXiv:2502.00989 [cs.CL] https://arxiv.org/abs/2502.00989

  91. [97]

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Ko...

  92. [98]

    Gower, Konstantin Korovin, Daniel Brunnsåker, Filip Kronström, Gabriel K

    Alexander H. Gower, Konstantin Korovin, Daniel Brunnsåker, Filip Kronström, Gabriel K. Reder, Ievgeniia A. Tiukova, Ronald S. Reiserer, John P. Wikswo, and Ross D. King. 2024. The Use of AI-Robotic Systems for Scientific Discovery. arXiv:2406.17835 [cs.LG] https://arxiv.org/ab...

  93. [99]

    Tianyang Gu, Jingjin Wang, Zhihao Zhang, and HaoHong Li. 2025. LLMs can Realize Combinatorial Creativity: Generating Creative Ideas via LLMs for Scientific Research. arXiv:2412.14141 [cs.AI] https://arxiv.org/abs/2412.14141

  94. [100]

    Yuzhe Gu, Wenwei Zhang, Chengqi Lyu, Dahua Lin, and Kai Chen. 2025. Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs. arXiv:2503.02846 [cs.CL] https://arxiv.org/abs/2503.02846

  95. [101]

    Chengquan Guo, Xun Liu, Chulin Xie, Andy Zhou, Yi Zeng, Zinan Lin, Dawn Song, and Bo Li. 2024. Red- Code: Risky Code Execution and Generation Benchmark for Code Agents. https://github.com/AI-secure/RedCode. arXiv:2411.07781 [cs.SE] https://arxiv.org/abs/2411.07781

  96. [102]

    Siyuan Guo, Cheng Deng, Ying Wen, Hechang Chen, Yi Chang, and Jun Wang. 2024. DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning. https://github.com/guosyjlu/DS-Agent. arXiv:2402.17453 [cs.LG] https://arxiv.org/abs/2402.17453

  97. [103]

    Xin Guo, Haotian Xia, Zhaowei Liu, Hanyang Cao, Zhi Yang, Zhiqiang Liu, Sizhe Wang, Jinyi Niu, Chuqi Wang, Yanhui Wang, Xiaolong Liang, Xiaoming Huang, Bing Zhu, Zhongyu Wei, Yun Chen, Weining Shen, and Liwen Zhang. 2024. FinEval: A Chinese Financial Domain Knowledge Evaluatio...

  98. [104]

    Hilda Hadan, Derrick Wang, Reza Hadi Mogavi, Joseph Tu, Leah Zhang-Kennedy, and Lennart E. Nacke. 2024. The Great AI Witch Hunt: Reviewers Perception and (Mis)Conception of Generative AI in Research Writing. https: //arxiv.org/abs/2407.12015

  99. [105]

    Sukjin Han. 2024. Mining Causality: AI-Assisted Search for Instrumental Variables. arXiv:2409.14202 [econ.EM] https://arxiv.org/abs/2409.14202

  100. [106]

    LaToza, and Brittany Johnson

    Ebtesam Al Haque, Chris Brown, Thomas D. LaToza, and Brittany Johnson. 2025. Towards Decoding Developer Cognition in the Age of AI Assistants. arXiv:2501.02684 [cs.HC] https://arxiv.org/abs/2501.02684

  101. [107]

    Gaole He, Patrick Hemmer, Michael Vössing, Max Schemmer, and Ujwal Gadiraju. 2025. Fine-Grained Appropriate Reliance: Human-AI Collaboration with a Multi-Step Transparent Decision Workflow for Complex Task Decomposition. arXiv:2501.10909 [cs.AI] https://arxiv.org/abs/2501.10909

  102. [108]

    Kaveen Hiniduma, Suren Byna, Jean Luca Bez, and Ravi Madduri. 2024. AI Data Readiness Inspector (AIDRIN) for Quantitative Assessment of Data Readiness for AI. InProceedings of the 36th International Conference on Scientific and Statistical Database Management (SSDBM 2024). ACM...

  103. [109]

    HKUDS. 2025. AI-Researcher. https://github.com/HKUDS/AI-Researcher

  104. [110]

    Brendan Hogan, Anmol Kabra, Felipe Siqueira Pacheco, Laura Greenstreet, Joshua Fan, Aaron Ferber, Marta Ummus, Alecsander Brito, Olivia Graham, Lillian Aoki, Drew Harvell, Alex Flecker, and Carla Gomes. 2024. AiSciVision: A Framework for Specializing Large Multimodal Models in...

  105. [111]

    Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024. MetaGPT: Meta Programming for A Multi-Agent Collaborative...

  106. [112]

    Hong Kong University Data Science Lab. 2024. Auto-Deep-Research. https://github.com/HKUDS/Auto-Deep- Research

  107. [113]

    Betty Li Hou, Kejian Shi, Jason Phang, James Aung, Steven Adler, and Rosie Campbell. 2024. Large Language Models as Misleading Assistants in Conversation. arXiv:2407.11789 [cs.CL] https://arxiv.org/abs/2407.11789

  108. [114]

    Shulin Huang, Shirong Ma, Yinghui Li, Mengzuo Huang, Wuhe Zou, Weidong Zhang, and Hai-Tao Zheng. 2024. LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles. https://github.com/THUKElab/LatEval. arXiv:2308.10855 [cs.CL] htt...

  109. [115]

    HuggingFace. 2025. smolagents: open_deep_research. https://github.com/huggingface/smolagents/tree/main/ examples/open_deep_research

  110. [116]

    Faria Huq, Abdus Samee, David Chuan-En Lin, Alice Xiaodi Tang, and Jeffrey P Bigham. 2025. NoTeeline: Supporting Real-Time, Personalized Notetaking with LLM-Enhanced Micronotes. InProceedings of the 30th International Conference on Intelligent User Interfaces (IUI ’25). ACM, 1...

  111. [117]

    Kurando IIDA and Kenjiro MIMURA. 2024. CATER: Leveraging LLM to Pioneer a Multidimensional, Reference- Independent Paradigm in Translation Quality Evaluation. arXiv:2412.11261 [cs.CL] https://arxiv.org/abs/2412.11261 86 Xu et al

  112. [118]

    Seyed Mohammad Ali Jafari. 2024. Streamlining the Selection Phase of Systematic Literature Reviews (SLRs) Using AI-Enabled GPT-4 Assistant API. arXiv:2402.18582 [cs.DL] https://arxiv.org/abs/2402.18582

  113. [119]

    Rishab Jain and Aditya Jain. 2023. Generative AI in Writing Research Papers: A New Type of Algorithmic Bias and Uncertainty in Scholarly Work. arXiv:2312.10057 [cs.CY] https://arxiv.org/abs/2312.10057

  114. [120]

    Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, and Yuxiang Wu. 2025. AIDE: AI-Driven Exploration in the Space of Code. arXiv:2502.13138 [cs.AI] https://arxiv.org/abs/2502.13138

  115. [121]

    Jina AI. 2025. node-DeepResearch. https://github.com/jina-ai/node-DeepResearch

  116. [122]

    Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu. 2025. DSBench: How Far Are Data Science Agents from Becoming Data Science Experts? https: //github.com/LiqiangJing/DSBench. arXiv:2409.07703 [cs.AI] https://arxi...

  117. [123]

    Nicola Jones. 2025. OpenAI’s ‘deep research’ tool: is it useful for scientists? https://www.nature.com/articles/d41586- 025-00377-9

  118. [124]

    Vijay Joshi and Iver Band. 2024. Disrupting Test Development with AI Assistants: Building the Base of the Test Pyramid with Three AI Coding Assistants. (Oct. 2024). doi:10.36227/techrxiv.173014488.82191966/v1

  119. [125]

    Majeed Kazemitabaar, Jack Williams, Ian Drosos, Tovi Grossman, Austin Zachary Henley, Carina Negreanu, and Advait Sarkar. 2024. Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition. In Proceedings of the 37th Annual ACM Symposium...

  120. [126]

    CTOL Editors Ken. 2025. Gemini Launches Deep Research on 2.5 Pro Aiming to Redefine AI-Powered Analysis with Strong Lead Over OpenAI. https://www.ctol.digital/news/gemini-deep-research-launch-2-5-pro-vs-openai/

  121. [127]

    Antti Keurulainen, Isak Westerlund, Samuel Kaski, and Alexander Ilin. 2021. Learning to Assist Agents by Observing Them. arXiv:2110.01311 [cs.AI] https://arxiv.org/abs/2110.01311

  122. [128]

    Abdullah Khalili and Abdelhamid Bouchachia. 2022. Toward Building Science Discovery Machines. arXiv:2103.15551 [cs.AI] https://arxiv.org/abs/2103.15551

  123. [129]

    Stefan Kramer, Mattia Cerrato, Sašo Džeroski, and Ross King. 2023. Automated Scientific Discovery: From Equation Discovery to Autonomous Discovery Systems. arXiv:2305.02251 [cs.AI] https://arxiv.org/abs/2305.02251

  124. [130]

    Kummerfeld, Anne Lauscher, Kevin Leyton-Brown, Sheng Lu, Mausam, Margot Mieskes, Aurélie Névéol, Danish Pruthi, Lizhen Qu, Roy Schwartz, Noah A

    Ilia Kuznetsov, Osama Mohammed Afzal, Koen Dercksen, Nils Dycke, Alexander Goldberg, Tom Hope, Dirk Hovy, Jonathan K. Kummerfeld, Anne Lauscher, Kevin Leyton-Brown, Sheng Lu, Mausam, Margot Mieskes, Aurélie Névéol, Danish Pruthi, Lizhen Qu, Roy Schwartz, Noah A. Smith, Thamar ...

  125. [131]

    Martin Lance. 2024. open_deep_research. https://github.com/langchain-ai/open_deep_research

  126. [132]

    Hao Lang, Fei Huang, and Yongbin Li. 2025. Debate Helps Weak-to-Strong Generalization. arXiv:2501.13124 [cs.CL] https://arxiv.org/abs/2501.13124

  127. [133]

    LangChain. 2025. How to think about agent frameworks. https://blog.langchain.dev/how-to-think-about-agent- frameworks/. https://docs.google.com/spreadsheets/d/1B37VxTBuGLeTSPVWtz7UMsCdtXrqV5hCjWkbHN8tfAo/

  128. [134]

    langChain AI. 2024. LangGraph. https://github.com/langchain-ai/langgraph

  129. [135]

    Andrew Laverick, Kristen Surrao, Inigo Zubeldia, Boris Bolliet, Miles Cranmer, Antony Lewis, Blake Sherwin, and Julien Lesgourgues. 2024. Multi-Agent System for Cosmological Parameter Analysis. arXiv:2412.00431 [astro-ph.IM] https://arxiv.org/abs/2412.00431

  130. [136]

    Eunhae Lee. 2024. Towards Ethical Personal AI Applications: Practical Considerations for AI Assistants with Long-Term Memory. arXiv:2409.11192 [cs.CY] https://arxiv.org/abs/2409.11192

  131. [137]

    Yuho Lee, Taewon Yun, Jason Cai, Hang Su, and Hwanjun Song. 2024. UniSumEval: Towards Unified, Fine- Grained, Multi-Dimensional Summarization Evaluation for LLMs. https://github.com/DISL-Lab/UniSumEval-v1.0. arXiv:2409.19898 [cs.CL] https://arxiv.org/abs/2409.19898

  132. [138]

    Letta-AI. 2023. Letta. https://github.com/letta-ai/letta

  133. [139]

    Berger, and Stephen N

    Kyla Levin, Nicolas van Kempen, Emery D. Berger, and Stephen N. Freund. 2025. ChatDBG: An AI-Powered Debugging Assistant. arXiv:2403.16354 [cs.SE] https://arxiv.org/abs/2403.16354

  134. [140]

    James R. Lewis. 2018. The System Usability Scale: Past, Present, and Future. International Journal of Human–Computer Interaction 34, 7 (2018), 577–590. doi:10.1080/10447318.2018.1455307 arXiv:https://doi.org/10.1080/10447318.2018.1455307

  135. [141]

    Li, Been Kim, and Zi Wang

    Belinda Z. Li, Been Kim, and Zi Wang. 2025. QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks? arXiv:2503.22674 [cs.AI] https://arxiv.org/abs/2503.22674

  136. [142]

    Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. arXiv:2303.17760 [cs.AI] https: //arxiv.org/abs/2303.17760 A Comprehensive Survey of Deep Resea...

  137. [143]

    Jiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube, Bingsheng Yao, Xuhai Xu, Dakuo Wang, Elizabeth Mynatt, and Varun Mishra. 2025. Vital Insight: Assisting Experts’ Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LL...

  138. [144]

    Xuefeng Li, Haoyang Zou, and Pengfei Liu. 2025. TORL: Scaling Tool-Integrated RL. https://github.com/GAIR- NLP/ToRL. https://arxiv.org/pdf/2503.23383

  139. [145]

    Yuan Li, Yixuan Zhang, and Lichao Sun. 2023. MetaAgents: Simulating Interactions of Human Behaviors for LLM-based Task-oriented Coordination via Collaborative Generative Agents. arXiv:2310.06500 [cs.AI] https://arxiv.org/abs/2310. 06500

  140. [146]

    Zhuoyan Li, Chen Liang, Jing Peng, and Ming Yin. 2024. How Does the Disclosure of AI Assistance Affect the Perceptions of Writing? arXiv:2410.04545 [cs.CL] https://arxiv.org/abs/2410.04545

  141. [147]

    Wilson, Woosang Lim, and William Yang Wang

    Zekun Li, Xianjun Yang, Kyuri Choi, Wanrong Zhu, Ryan Hsieh, HyeonJung Kim, Jin Hyuk Lim, Sungyoung Ji, Byungju Lee, Xifeng Yan, Linda Ruth Petzold, Stephen D. Wilson, Woosang Lim, and William Yang Wang. 2025. MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Sci...

  142. [148]

    Liang, Chenyang Yang, and Brad A

    Jenny T. Liang, Chenyang Yang, and Brad A. Myers. 2023. A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and Challenges. arXiv:2303.17125 [cs.SE] https://arxiv.org/abs/2303.17125

  143. [149]

    Manning, Christopher Ré, Diana Acosta-Navas, Drew A

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...

  144. [150]

    Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and Xiaodong Shi. 2023. Automated scholarly paper review: Concepts, technologies, and challenges.Information Fusion98 (Oct. 2023), 101830. doi:10.1016/j.inffus.2023.101830

  145. [151]

    Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv:2109.07958 [cs.CL] https://arxiv.org/abs/2109.07958

  146. [152]

    Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, Jianfeng Gao, and Chunyuan Li. 2023. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents. arXiv:2311.05437 [cs.CV] https://arxiv.org/abs/2311.05437

  147. [154]

    Xiao Liu, Bo Qin, Dongzhu Liang, Guang Dong, Hanyu Lai, Hanchen Zhang, Hanlin Zhao, Iat Long Iong, Jiadai Sun, Jiaqi Wang, Junjie Gao, Junjun Shan, Kangning Liu, Shudan Zhang, Shuntian Yao, Siyi Cheng, Wentao Yao, Wenyi Zhao, Xinghan Liu, Xinyi Liu, Xinying Chen, Xinyue Yang, ...

  148. [155]

    Zijun Liu, Kaiming Liu, Yiqi Zhu, Xuanyu Lei, Zonghan Yang, Zhenhe Zhang, Peng Li, and Yang Liu. 2024. AIGS: Generating Science from AI-Powered Automated Falsification. arXiv:2411.11910 [cs.LG] https://arxiv.org/abs/2411. 11910

  149. [156]

    Zhiwei Liu, Weiran Yao, Jianguo Zhang, Le Xue, Shelby Heinecke, Rithesh Murthy, Yihao Feng, Zeyuan Chen, Juan Carlos Niebles, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese. 2023. BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Ag...

  150. [157]

    Renze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Jian Xie, Yuxuan Sun, Yusen Zhang, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Wenchao Ma, Xi Li, Kai Zhang, Congying Xia, Lifu Huang, and Wenpeng Yin. 2025. AAAR-1.0: Assessing AI’s Potential to Assist R...

  151. [158]

    Cong Lu, Shengran Hu, and Jeff Clune. 2025. Automated Capability Discovery via Model Self-Exploration. arXiv:2502.07577 [cs.LG] https://arxiv.org/abs/2502.07577 88 Xu et al

  152. [159]

    Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292 [cs.AI] https://arxiv.org/abs/2408.06292

  153. [160]

    Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. https://scienceqa.github.io/. arXiv:2209.09513 [cs.CL] htt...

  154. [161]

    Chandra Maddila, Negar Ghorbani, Kosay Jabre, Vijayaraghavan Murali, Edwin Kim, Parth Thakkar, Nikolay Pavlovich Laptev, Olivia Harman, Diana Hsu, Rui Abreu, and Peter C. Rigby. 2024. AI-Assisted SQL Authoring at Industry Scale. arXiv:2407.13280 [cs.SE] https://arxiv.org/abs/2...

  155. [162]

    Srijoni Majumdar, Edith Elkind, and Evangelos Pournaras. 2025. Generative AI Voting: Fair Collective Choice is Resilient to LLM Biases and Inconsistencies. arXiv:2406.11871 [cs.AI] https://arxiv.org/abs/2406.11871

  156. [163]

    Doan,Nam V.Nguyen, QuangPham, andNghiD

    DungNguyenManh, ThangPhanChau, NamLeHai, ThongT. Doan,Nam V.Nguyen, QuangPham, andNghiD. Q.Bui

  157. [164]

    Manus. 2025. Manus. https://manus.im/

  158. [165]

    Rohin Manvi, Samar Khanna, Gengchen Mai, Marshall Burke, David Lobell, and Stefano Ermon. 2024. GeoLLM: Extracting Geospatial Knowledge from Large Language Models. arXiv:2310.06213 [cs.CL] https://arxiv.org/abs/2310. 06213

  159. [166]

    Markowitz

    David M. Markowitz. 2024. From Complexity to Clarity: How AI Enhances Perceptions of Scientists and the Public’s Understanding of Science. arXiv:2405.00706 [cs.CL] https://arxiv.org/abs/2405.00706

  160. [167]

    Jonathan Mast. 2025. ChatGPT’s Deep Research vs. Google’s Gemini 1.5 Pro with Deep Research: A Detailed Comparison. https://whitebeardstrategies.com/ai-prompt-engineering/chatgpts-deep-research-vs-googles-gemini-1-5- pro-with-deep-research-a-detailed-comparison/

  161. [168]

    Mastra-AI. 2025. Mastra. https://github.com/mastra-ai/mastra

  162. [169]

    Shray Mathur, Noah van der Vleuten, Kevin Yager, and Esther Tsai. 2024. VISION: A Modular AI Assistant for Natural Human-Instrument Interaction at Scientific User Facilities. arXiv:2412.18161 [cs.AI] https://arxiv.org/abs/2412.18161

  163. [170]

    Gianmarco Mengaldo. 2025. Explain the Black Box for the Sake of Science: the Scientific Method in the Era of Generative Artificial Intelligence. arXiv:2406.10557 [cs.AI] https://arxiv.org/abs/2406.10557

  164. [171]

    2025.MGX.dev

    MGX Technologies. 2025.MGX.dev. https://mgx.dev

  165. [172]

    Gregoire Mialon, Clementine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023. GAIA:A Benchmark for General AI Assistants. https://huggingface.co/gaia-benchmark. https://arxiv.org/pdf/2311.12983

  166. [173]

    Microsoft. 2023. Microsoft Copilot. https://www.microsoft.com/en-us/microsoft-copilot/organizations

  167. [174]

    Microsoft. 2023. Semantic-kernel. https://github.com/microsoft/semantic-kernel

  168. [175]

    mirayayerdem. 2022. Github-Copilot-Amazon-Whisperer-ChatGPT. https://github.com/mirayayerdem/Github-Copilot- Amazon-Whisperer-ChatGPT

  169. [176]

    Mlc-ai. 2023. web-llm. https://github.com/mlc-ai/web-llm

  170. [177]

    ModelTC. 2025. lightllm. https://github.com/ModelTC/lightllm

  171. [178]

    Devam Mondal and Atharva Inamdar. 2024. SeqMate: A Novel Large Language Model Pipeline for Automating RNA Sequencing. arXiv:2407.03381 [q-bio.GN] https://arxiv.org/abs/2407.03381

  172. [179]

    Peya Mowar, Yi-Hao Peng, Jason Wu, Aaron Steinfeld, and Jeffrey P. Bigham. 2025. CodeA11y: Making AI Coding Assistants Useful for Accessible Web Development. arXiv:2502.10884 [cs.HC] https://arxiv.org/abs/2502.10884

  173. [180]

    Dilxat Muhtar, Zhenshi Li, Feng Gu, Xueliang Zhang, and Pengfeng Xiao. 2024. LHRS-Bot: Empowering Re- mote Sensing with VGI-Enhanced Large Multimodal Language Model. https://github.com/NJU-LHRS/LHRS-Bot. arXiv:2402.02544 [cs.CV] https://arxiv.org/abs/2402.02544

  174. [181]

    Manisha Mukherjee, Sungchul Kim, Xiang Chen, Dan Luo, Tong Yu, and Tung Mai. 2025. From Documents to Dialogue: Building KG-RAG Enhanced AI Assistants. arXiv:2502.15237 [cs.IR] https://arxiv.org/abs/2502.15237

  175. [182]

    Sheshera Mysore, Mahmood Jasim, Haoru Song, Sarah Akbar, Andre Kenneth Chase Randall, and Narges Mahyar

  176. [183]

    n8n. 2023. n8n. https://github.com/n8n-io/n8n

  177. [184]

    Nanobrowser Team. 2024. Nanobrowser. https://github.com/nanobrowser/nanobrowser

  178. [185]

    Nathalia Nascimento, Everton Guimaraes, Sai Sanjna Chintakunta, and Santhosh Anitha Boominathan. 2024. LLM4DS: Evaluating Large Language Models for Data Science Code Generation. https://github.com/DataForScience/LLM4DS. arXiv:2411.11908 [cs.SE] https://arxiv.org/abs/2411.11908

  179. [186]

    InProceedings of the 2023 Conference on Human Information Interaction and Retrieval (CHIIR ’23)

    How Data Scientists Review the Scholarly Literature. InProceedings of the 2023 Conference on Human Information Interaction and Retrieval (CHIIR ’23). ACM, 137–152. doi:10.1145/3576840.3578309

  180. [187]

    Alex Nguyen, Zilong Wang, Jingbo Shang, and Dheeraj Mekala. 2024. DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering. arXiv:2404.00439 [cs.CL] https://arxiv.org/abs/2404.00439

  181. [188]

    Nguyen, Fengchun Qiao, Arthur Trembanis, and Xi Peng

    Kien X. Nguyen, Fengchun Qiao, Arthur Trembanis, and Xi Peng. 2024. SeafloorAI: A Large-scale Vision-Language Dataset for Seafloor Geological Survey. https://github.com/deep-real/SeafloorAI. arXiv:2411.00172 [cs.CV] https: //arxiv.org/abs/2411.00172

  182. [189]

    Ziqi Ni, Yahao Li, Kaijia Hu, Kunyuan Han, Ming Xu, Xingyu Chen, Fengqi Liu, Yicong Ye, and Shuxin Bai

  183. [190]

    Khanh Nghiem, Anh Minh Nguyen, and Nghi D. Q. Bui. 2024. Envisioning the Next-Generation AI Coding Assistants: Insights & Proposals. arXiv:2403.14592 [cs.SE] https://arxiv.org/abs/2403.14592 A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications 89

  184. [191]

    Kanda, and Haruka Ozaki

    Koji Ochiai, Yuya Tahara-Arai, Akari Kato, Kazunari Kaizu, Hirokazu Kariyazaki, Makoto Umeno, Koichi Takahashi, Genki N. Kanda, and Haruka Ozaki. 2025. Automating Care by Self-maintainability for Full Laboratory Automation. arXiv:2501.05789 [q-bio.QM] https://arxiv.org/abs/2501.05789

  185. [192]

    Ollama. 2023. Ollama. https://github.com/ollama/ollama

  186. [193]

    Open Manus Team. 2025. OpenManus. https://github.com/mannaandpoem/OpenManus

  187. [194]

    arXiv:2411.08063 [physics.soc-ph] https://arxiv.org/abs/2411.08063

    MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration. arXiv:2411.08063 [physics.soc-ph] https://arxiv.org/abs/2411.08063

  188. [195]

    Alexander Novikov, Ngân V˜ u, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaud- huri, George Holland, Alex Davies, Sebastian Nowozin, Pu...

  189. [196]

    OpenAI. 2025. Deep Research System Card. https://cdn.openai.com/deep-research-system-card.pdf

  190. [198]

    OpenAI. 2025. Introducing OpenAI o3 and o4-mini. https://openai.com/index/introducing-o3-and-o4-mini/

  191. [199]

    OpenAI. 2025. codex. https://github.com/openai/codex

  192. [200]

    OpenAI. 2025. Compare models - OpenAI API. https://platform.openai.com/docs/models/compare?model=o3

  193. [201]

    OpenAI. 2025. Thinking with images. https://openai.com/index/thinking-with-images/

  194. [202]

    OpenBMB. 2023. XAgent. https://github.com/OpenBMB/XAgent

  195. [203]

    Orkes. 2022. Orkes. https://orkes.io/use-cases/agentic-workflows

  196. [204]

    OpenAI. 2025. OpenAI Agents SDK. https://github.com/openai/openai-agents-python

  197. [205]

    OpenAI. 2025. OpenAI o3 and o4-mini System Card. https://cdn.openai.com/pdf/2221c875-02dc-4789-800b- e7758f3722c1/o3-and-o4-mini-system-card.pdf

  198. [206]

    Mahabubur Rahman, and Mst

    Md Sultanul Islam Ovi, Nafisa Anjum, Tasmina Haque Bithe, Md. Mahabubur Rahman, and Mst. Shahnaj Akter Smrity

  199. [207]

    Carlos Alves Pereira, Tanay Komarlu, and Wael Mobeirek. 2023. The Future of AI-Assisted Writing. arXiv:2306.16641 [cs.HC] https://arxiv.org/abs/2306.16641

  200. [208]

    Mike Perkins and Jasper Roe. 2024. Generative AI Tools in Academic Research: Applications and Implications for Qualitative and Quantitative Research Methodologies. arXiv:2408.06872 [cs.HC] https://arxiv.org/abs/2408.06872

  201. [209]

    Takauki Osogami. 2025. Position: AI agents should be regulated based on autonomous action sequences. arXiv:2503.04750 [cs.CY] https://arxiv.org/abs/2503.04750

  202. [210]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...

  203. [211]

    Tomas Petricek, Gerrit J. J. van den Burg, Alfredo Nazábal, Taha Ceritli, Ernesto Jiménez-Ruiz, and Christopher K. I. Williams. 2022. AI Assistants: A Framework for Semi-Automated Data Wrangling. arXiv:2211.00192 [cs.DB] https://arxiv.org/abs/2211.00192

  204. [212]

    arXiv:2409.19922 [cs.SE] https://arxiv.org/abs/2409.19922

    Benchmarking ChatGPT, Codeium, and GitHub Copilot: A Comparative Study of AI-Driven Programming and Debugging Assistants. arXiv:2409.19922 [cs.SE] https://arxiv.org/abs/2409.19922

  205. [213]

    Evangelos Pournaras. 2023. Science in the Era of ChatGPT, Large Language Models and Generative AI: Challenges for Research Ethics and How to Respond. arXiv:2305.15299 [cs.CY] https://arxiv.org/abs/2305.15299

  206. [214]

    Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024. Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented 90 Xu et al. Generation Track. arXiv:2406.16828 [cs.IR] https://...

  207. [215]

    Perplexity. 2025. Introducing Perplexity Deep Research. https://www.perplexity.ai/hub/blog/introducing-perplexity- deep-research

  208. [216]

    Perplexity. 2025. Sonar by Perplexity. https://docs.perplexity.ai/guides/model-cards#research-models

  209. [217]

    Pythagora-io. 2024. gpt-pilot. https://github.com/Pythagora-io/gpt-pilot

  210. [218]

    Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. 2025. Humanity’s Last Exam. arXiv:2501.14249 [cs.LG] https://arxiv.org/abs/ 2501.14249

  211. [219]

    Laryn Qi, J. D. Zamfirescu-Pereira, Taehan Kim, Björn Hartmann, John DeNero, and Narges Norouzi. 2024. A Knowledge-Component-Based Methodology for Evaluating AI Assistants. arXiv:2406.05603 [cs.CY] https://arxiv.org/ abs/2406.05603

  212. [220]

    Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024. ChatDev: Communicative Agents for Software Development. https://github.com/OpenBMB/ChatDev. https://aclant...

  213. [221]

    Reeves, Jaromir Savelka, David H

    James Prather, Juho Leinonen, Natalie Kiesler, Jamie Gorson Benario, Sam Lau, Stephen MacNeil, Narges Norouzi, Simone Opel, Vee Pettit, Leo Porter, Brent N. Reeves, Jaromir Savelka, David H. Smith IV, Sven Strickroth, and Daniel Zingaro. 2024. Beyond the Hype: A Comprehensive ...

  214. [222]

    Pydantic. 2024. Pydantic-AI. https://github.com/pydantic/pydantic-ai

  215. [223]

    Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023. ToolLLM: Facilitating Large Language Model...

  216. [224]

    Jingyuan Qi, Zian Jia, Minqian Liu, Wangzhi Zhan, Junkai Zhang, Xiaofei Wen, Jingru Gan, Jianpeng Chen, Qin Liu, Mingyu Derek Ma, Bangzheng Li, Haohui Wang, Adithya Kulkarni, Muhao Chen, Dawei Zhou, Ling Li, Wei Wang, and Lifu Huang. 2024. MetaScientist: A Human-AI Synergistic...

  217. [225]

    Joaquin Ramirez-Medina, Mohammadmehdi Ataei, and Alidad Amirfazli. 2025. Accelerating Scientific Research Through a Multi-LLM Framework. arXiv:2502.07960 [physics.app-ph] https://arxiv.org/abs/2502.07960

  218. [226]

    Ruchit Rawal, Victor-Alexandru Pădurean, Sven Apel, Adish Singla, and Mariya Toneva. 2024. Hints Help Finding and Fixing Bugs Differently in Python and Text-based Program Representations. arXiv:2412.12471 [cs.SE] https: //arxiv.org/abs/2412.12471

  219. [227]

    Shuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang, Ningyu Zhang, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen. 2025. Benchmarking Agentic Workflow Generation. https://github.com/zjunlp/WorfBench. arXiv:2410.07869 [cs.CL] https://arxiv.org/abs/2410.07869

  220. [229]

    Restate. 2024. Restate. https://restate.dev/

  221. [230]

    Qwen LM. 2024. Qwen-Agent. https://github.com/QwenLM/Qwen-Agent

  222. [231]

    Filippo Ricca, Alessandro Marchetto, and Andrea Stocco. 2025. A Multi-Year Grey Literature Review on AI-assisted Test Automation. https://arxiv.org/pdf/2408.06224

  223. [232]

    Nathalie Riche, Anna Offenwanger, Frederic Gmeiner, David Brown, Hugo Romat, Michel Pahud, Nicolai Marquardt, Kori Inkpen, and Ken Hinckley. 2025. AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools. http...

  224. [233]

    Runtao Ren, Jian Ma, and Jianxi Luo. 2025. Large language model for patent concept generation.Advanced Engineering Informatics 65 (May 2025), 103301. doi:10.1016/j.aei.2025.103301

  225. [234]

    ResearchRabbit. 2025. ResearchRabbit. https://www.researchrabbit.ai/

  226. [235]

    Run-llama. 2023. LlamaIndex. https://github.com/run-llama/llama_index

  227. [236]

    reworkd. 2023. AgentGPT. https://github.com/reworkd/AgentGPT

  228. [237]

    SamuelSchmidgall. 2025. AgentLaboratory. https://github.com/SamuelSchmidgall/AgentLaboratory. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications 91

  229. [238]

    Huberman

    Thomas Sandholm, Sarah Dong, Sayandev Mukherjee, John Feland, and Bernardo A. Huberman. 2024. Semantic Navigation for AI-assisted Ideation. arXiv:2411.03575 [cs.HC] https://arxiv.org/abs/2411.03575

  230. [239]

    Lavista Ferres

    Anthony Cintron Roman, Jennifer Wortman Vaughan, Valerie See, Steph Ballard, Jehu Torres, Caleb Robinson, and Juan M. Lavista Ferres. 2024. Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments. arXiv:2312.06153 [cs.LG] https://arxiv....

  231. [240]

    Kaushik Roy, Vedant Khandelwal, Harshul Surana, Valerie Vera, Amit Sheth, and Heather Heckman. 2023. GEAR-Up: Generative AI and External Knowledge-based Retrieval Upgrading Scholarly Article Searches for Systematic Reviews. arXiv:2312.09948 [cs.IR] https://arxiv.org/abs/2312.09948

  232. [241]

    Schuemie, M

    Martijn J. Schuemie, M. Soledad Cepeda, Marc A. Suchard, Jianxiao Yang, Yuxi Tian, Alejandro Schuler, Patrick B. Ryan, David Madigan, and George Hripcsak. 2020. How Confident Are We About Observational Findings in Healthcare: A Benchmark Study.Harvard Data Science Review2, 1 (...

  233. [242]

    Sergey V Samsonau, Aziza Kurbonova, Lu Jiang, Hazem Lashen, Jiamu Bai, Theresa Merchant, Ruoxi Wang, Laiba Mehnaz, Zecheng Wang, and Ishita Patil. 2024. Artificial Intelligence for Scientific Research: Authentic Research Education Framework. arXiv:2210.08966 [cs.CY] https://ar...

  234. [243]

    Scite. 2025. Scite. https://scite.ai/

  235. [244]

    Agnia Sergeyuk, Yaroslav Golubev, Timofey Bryksin, and Iftekhar Ahmed. 2025. Using AI-based coding assistants in practice: State of affairs, perceptions, and ways forward.Information and Software Technology178 (Feb. 2025), 107610. doi:10.1016/j.infsof.2024.107610

  236. [245]

    Lindsay Sanneman and Julie Shah. 2021. Explaining Reward Functions to Humans for Better Human-Robot Collabo- ration. arXiv:2110.04192 [cs.RO] https://arxiv.org/abs/2110.04192

  237. [246]

    Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum. 2025. Agent Laboratory: Using LLM Agents as Research Assistants. arXiv:2501.04227 [cs.HC] https://arxiv.org/abs/2501.04227

  238. [247]

    Zejiang Shen, Tal August, Pao Siangliulue, Kyle Lo, Jonathan Bragg, Jeff Hammerbacher, Doug Downey, Joseph Chee Chang, and David Sontag. 2023. Beyond Summarization: Designing AI Support for Real-World Expository Writing Tasks. arXiv:2304.02623 [cs.CL] https://arxiv.org/abs/2304.02623

  239. [248]

    Scispace. 2024. Scispace. https://scispace.com/

  240. [249]

    Michael Shumer. 2025. OpenDeepResearcher. https://github.com/mshumer/OpenDeepResearcher

  241. [250]

    Significant-Gravitas. 2023. AutoGPT. https://github.com/Significant-Gravitas/AutoGPT

  242. [251]

    Mahsa Shamsabadi and Jennifer D’Souza. 2024. A FAIR and Free Prompt-based Research Assistant. arXiv:2405.14601 [cs.CL] https://arxiv.org/abs/2405.14601

  243. [252]

    Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. https://github.com/microsoft/JARVIS. https: //arxiv.org/pdf/2303.17580

  244. [253]

    Michael Skarlinski, Tyler Nadolski, James Braza, Remo Storni, Mayk Caldas, Ludovico Mitchener, Michaela Hinks, Andrew White, and Sam Rodriques. 2025. FutureHouse Platform: Superintelligent AI Agents for Scientific Discovery. https://www.futurehouse.org/research-announcements/l...

  245. [254]

    Shuming Shi, Enbo Zhao, Duyu Tang, Yan Wang, Piji Li, Wei Bi, Haiyun Jiang, Guoping Huang, Leyang Cui, Xinting Huang, Cong Zhou, Yong Dai, and Dongyang Ma. 2022. Effidit: Your AI Writing Assistant. arXiv:2208.01815 [cs.CL] https://arxiv.org/abs/2208.01815

  246. [255]

    Jamshid Sourati and James Evans. 2021. Accelerating science with human versus alien artificial intelligences. arXiv:2104.05188 [cs.AI] https://arxiv.org/abs/2104.05188

  247. [256]

    Jamshid Sourati and James Evans. 2023. Accelerating science with human-aware artificial intelligence. arXiv:2306.01495 [cs.AI] https://arxiv.org/abs/2306.01495

  248. [257]

    David Silver and Richard Sutton. 2025. Welcome to the Era of Experience. https://storage.googleapis.com/deepmind- media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf

  249. [258]

    It is there, and you need it, so why do you not use it?

    Auste Simkute, Ewa Luger, Michael Evans, and Rhianne Jones. 2024. "It is there, and you need it, so why do you not use it?" Achieving better adoption of AI systems by domain experts, in the case study of natural science research. arXiv:2403.16895 [cs.HC] https://arxiv.org/abs/...

  250. [259]

    Suzhou Yuling Artificial Intelligence Technology Co

    Ltd. Suzhou Yuling Artificial Intelligence Technology Co. 2023. Dify: Open-source LLM Application Development Platform. https://dify.ai/

  251. [260]

    Clark, Hao He, Haoran He, Jie Min, Xinlei Zhang, Simin Zheng, Zhiyang Zhang, Xinwei Deng, and Yili Hong

    Xinyi Song, Kexin Xie, Lina Lee, Ruizhe Chen, Jared M. Clark, Hao He, Haoran He, Jie Min, Xinlei Zhang, Simin Zheng, Zhiyang Zhang, Xinwei Deng, and Yili Hong. 2025. Performance Evaluation of Large Language Models in Statistical Programming. arXiv:2502.13117 [stat.AP] https://...

  252. [261]

    Brian Tang and Kang G. Shin. 2024. Steward: Natural Language Web Automation. arXiv:2409.15441 [cs.AI] https://arxiv.org/abs/2409.15441

  253. [262]

    Jiabin Tang, Tianyu Fan, and Chao Huang. 2025. AutoAgent: A Fully-Automated and Zero-Code Framework for LLM Agents. arXiv:2502.05957 [cs.AI] https://arxiv.org/abs/2502.05957

  254. [263]

    StanfordNLP. 2024. DSPy. https://github.com/stanfordnlp/dspy

  255. [264]

    Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, and Nanqing Dong. 2025. Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System. arXiv:24...

  256. [265]

    Temporalio. 2020. Temporal. https://github.com/temporalio/temporal

  257. [266]

    Xin Tan, Xiao Long, Xianjun Ni, Yinghao Zhu, Jing Jiang, and Li Zhang. 2024. How far are AI-powered programming assistants from meeting developers’ needs? arXiv:2404.12000 [cs.SE] https://arxiv.org/abs/2404.12000

  258. [267]

    TheBlewish. 2024. Automated-AI-Web-Researcher-Ollama. https://github.com/TheBlewish/Automated-AI-Web- Researcher-Ollama

  259. [268]

    Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, Shengyan Liu, Di Luo, Yutao Ma, Hao Tong, Kha Trinh, Chenyu Tian, Zihan Wang, Bohao Wu, Yanyu Xiong, Shengzhu Yin, Minhui Zhu, Kilian Lieret, Yanx...

  260. [269]

    Yan Tang. 2025. deep_research_agent. https://github.com/grapeot/deep_research_agent. 92 Xu et al

  261. [270]

    Tadahiro Taniguchi, Shiro Takagi, Jun Otsuka, Yusuke Hayashi, and Hiro Taiyo Hamada. 2024. Collective Predictive Coding as Model of Science: Formalizing Scientific Activities Towards Generative Science. arXiv:2409.00102 [physics.soc- ph] https://arxiv.org/abs/2409.00102

  262. [271]

    Benjamin Towle and Ke Zhou. 2024. Enhancing AI Assisted Writing with One-Shot Implicit Negative Feedback. arXiv:2410.11009 [cs.CL] https://arxiv.org/abs/2410.11009

  263. [272]

    Enkeleda Thaqi, Mohamed Omar Mantawy, and Enkelejda Kasneci. 2024. SARA: Smart AI Reading Assistant for Reading Comprehension. InProceedings of the 2024 Symposium on Eye Tracking Research and Applications (ETRA ’24). ACM, 1–3. doi:10.1145/3649902.3655661

  264. [273]

    Wang, Sabrina A Sgandurra, Reza Hadi Mogavi, and Lennart E

    Joseph Tu, Hilda Hadan, Derrick M. Wang, Sabrina A Sgandurra, Reza Hadi Mogavi, and Lennart E. Nacke. 2024. Augmenting the Author: Exploring the Potential of AI Collaboration in Academic Writing. arXiv:2404.16071 [cs.HC] https://arxiv.org/abs/2404.16071

  265. [274]

    Su, and Linjun Zhang

    Xinming Tu, James Zou, Weijie J. Su, and Linjun Zhang. 2023. What Should Data Science Education Do with Large Language Models? arXiv:2307.02792 [cs.CY] https://arxiv.org/abs/2307.02792

  266. [275]

    Tiukova, Daniel Brunnsåker, Erik Y

    Ievgeniia A. Tiukova, Daniel Brunnsåker, Erik Y. Bjurström, Alexander H. Gower, Filip Kronström, Gabriel K. Reder, Ronald S. Reiserer, Konstantin Korovin, Larisa B. Soldatova, John P. Wikswo, and Ross D. King. 2024. Genesis: Towards the Automation of Systems Biology Research. ...

  267. [276]

    Irina Tolstykh, Aleksandra Tsybina, Sergey Yakubson, Aleksandr Gordeev, Vladimir Dokholyan, and Maksim Kuprashe- vich. 2024. GigaCheck: Detecting LLM-generated Content. arXiv:2410.23728 [cs.CL] https://arxiv.org/abs/2410.23728

  268. [277]

    Rasmus Ulfsnes, Nils Brede Moe, Viktoria Stray, and Marianne Skarpen. 2024. Transforming Software Development with Generative AI: Empirical Insights on Collaboration and Workflow. arXiv:2405.01543 [cs.SE] https://arxiv.org/ abs/2405.01543

  269. [278]

    Thanh-Dat Truong, Hoang-Quan Nguyen, Xuan-Bac Nguyen, Ashley Dowling, Xin Li, and Khoa Luu. 2025. Insect- Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding. https: //uark-cviu.github.io/projects/insect-foundation/. arXiv:2502....

  270. [279]

    Jones, Oisin Mac Aodha, Sara Beery, and Grant Van Horn

    Edward Vendrow, Omiros Pantazis, Alexander Shepard, Gabriel Brostow, Kate E. Jones, Oisin Mac Aodha, Sara Beery, and Grant Van Horn. 2024. INQUIRE: A Natural World Text-to-Image Retrieval Benchmark. https://inquire- benchmark.github.io/. arXiv:2411.02537 [cs.CV] https://arxiv....

  271. [280]

    Vercel. 2020. Vercel. https://vercel.com/

  272. [281]

    Michele Tufano, Anisha Agarwal, Jinu Jang, Roshanak Zilouchian Moghaddam, and Neel Sundaresan. 2024. AutoDev: Automated AI-Driven Development. arXiv:2403.08299 [cs.SE] https://arxiv.org/abs/2403.08299

  273. [282]

    Aleksei Turobov, Diane Coyle, and Verity Harding. 2024. Using ChatGPT for Thematic Analysis. arXiv:2405.08828 [cs.HC] https://arxiv.org/abs/2405.08828

  274. [283]

    Weisz, Xuye Liu, Lingfei Wu, and Casey Dugan

    April Yi Wang, Dakuo Wang, Jaimie Drozdal, Michael Muller, Soya Park, Justin D. Weisz, Xuye Liu, Lingfei Wu, and Casey Dugan. 2022. Documentation Matters: Human-Centered AI System to Assist Data Science Code Documentation in Computational Notebooks.ACM Transactions on Computer...

  275. [284]

    Stanford University. 2025. STORM. https://storm.genie.stanford.edu/

  276. [285]

    Suyuan Wang, Xueqian Yin, Menghao Wang, Ruofeng Guo, and Kai Nan. 2024. EvoPat: A Multi-LLM-based Patents Summarization and Analysis Agent. arXiv:2412.18100 [cs.DL] https://arxiv.org/abs/2412.18100

  277. [286]

    Tiannan Wang, Jiamin Chen, Qingrui Jia, Shuai Wang, Ruoyu Fang, Huilin Wang, Zhaowei Gao, Chunzhao Xie, Chuou Xu, Jihong Dai, Yibin Liu, Jialong Wu, Shengwei Ding, Long Li, Zhiwei Huang, Xinle Deng, Teng Yu, Gangan Ma, Han A Comprehensive Survey of Deep Research: Systems, Meth...

  278. [287]

    Vllm-project. 2023. vllm. https://github.com/vllm-project/vllm

  279. [288]

    Thiemo Wambsganss, Xiaotian Su, Vinitra Swamy, Seyed Parsa Neshaei, Roman Rietsche, and Tanja Käser. 2023. Unraveling Downstream Gender Bias from Large Language Models: A Study on AI Educational Writing Assistance. arXiv:2311.03311 [cs.CL] https://arxiv.org/abs/2311.03311

  280. [289]

    Ying-Mei Wang and Tzeng-J Chen. 2025. AI’s deep research revolution: Transforming biomedical literature analysis. https://journals.lww.com/jcma/citation/9900/ai_s_deep_research_revolution__transforming.508.aspx

  281. [290]

    Vera Liao, Yunfeng Zhang, Udayan Khurana, Horst Samulowitz, Soya Park, Michael Muller, and Lisa Amini

    Dakuo Wang, Q. Vera Liao, Yunfeng Zhang, Udayan Khurana, Horst Samulowitz, Soya Park, Michael Muller, and Lisa Amini. 2021. How Much Automation Does a Data Scientist Want? arXiv:2101.03970 [cs.LG] https://arxiv.org/abs/ 2101.03970

  282. [291]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903 [cs.CL] https://arxiv.org/abs/2201.11903

  283. [292]

    Shufa Wei, Xiaolong Xu, Xianbiao Qi, Xi Yin, Jun Xia, Jingyi Ren, Peijun Tang, Yuxiang Zhong, Yihao Chen, Xiaoqin Ren, Yuxin Liang, Liankai Huang, Kai Xie, Weikang Gui, Wei Tan, Shuanglong Sun, Yongquan Hu, Qinxian Liu, Nanjin Li, Chihao Dai, Lihua Wang, Xiaohui Liu, Lei Zhang...

  284. [293]

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171 [cs.CL] https://arxiv.org/abs/2203.11171

  285. [294]

    Yao Wang, Mingxuan Cui, and Arthur Jiang. 2025. Enabling AI Scientists to Recognize Innovation: A Domain-Agnostic Algorithm for Assessing Novelty. arXiv:2503.01508 [cs.AI] https://arxiv.org/abs/2503.01508

  286. [296]

    Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024. Measuring short-form factuality in large language models. https://cdn.openai.com/papers/ simpleqa.pdf

  287. [300]

    Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. 2025. CycleResearcher: Improving Automated Research via Automated Review. arXiv:2411.00816 [cs.CL] https://arxiv.org/ abs/2411.00816

  288. [2023]

    arXiv:2307.07049 [cs.CL] https://arxiv.org/abs/2307.07049 82 Xu et al

    MegaWika: Millions of reports and their sources across 50 diverse languages. arXiv:2307.07049 [cs.CL] https://arxiv.org/abs/2307.07049 82 Xu et al

  289. [2024]

    https: //uni-medical.github.io/GMAI-MMBench.github.io/

    GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI. https: //uni-medical.github.io/GMAI-MMBench.github.io/. arXiv:2408.03361 [eess.IV] https://arxiv.org/abs/2408.03361

  290. [2025]

    https://github.com/FSoft-AI4Code/CodeMMLU

    CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs. https://github.com/FSoft-AI4Code/CodeMMLU. arXiv:2410.01999 [cs.SE] https://arxiv.org/abs/2410.01999

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.