REVIEW 4 major objections 5 minor 12 cited by
A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A new four-way taxonomy organizes the AI 'Deep Research' boom, and its feature tables rest on vendor-reported numbers that were only partially verified.
desk verdict A useful organizational survey of the Deep Research ecosystem, undermined by contradictory benchmark tables that need reconciliation before the comparative claims can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the four-dimension hierarchical taxonomy itself: foundation models and reasoning engines, tool utilization and environmental interaction, task planning and execution control, and knowledge synthesis and output generation. The taxonomy organizes everything else in the survey, and it is paired with a four-pattern architectural analysis (monolithic, pipeline, multi-agent, hybrid) that explains how systems manage control flow, component coupling, failure propagation, and deployment flexibility.
What would settle it
A controlled evaluation of several systems (e.g., OpenAI/DeepResearch, Gemini/DeepResearch, Perplexity/DeepResearch, and one open-source alternative) on identical tasks from HLE and GAIA, run under the same protocol with repeated trials, would reveal whether the reported margins hold or whether the benchmark-based ranking is an artifact of vendor self-reporting.
Extended reading notes
Core claim
The central claim is that the diverse landscape of Deep Research systems can be productively organized by a four-dimensional hierarchical taxonomy. Using that taxonomy as a lens, the paper maps the evolution from general-purpose LLM assistants to specialized research systems, identifies four recurring architectural patterns (monolithic, pipeline-based, multi-agent, and hybrid), compares representative systems on benchmarks such as HLE, MMLU, HotpotQA, and GAIA, and evaluates their suitability across academic, enterprise, financial, educational, and personal knowledge-management applications. The intended contribution is both theoretical, a framework for comparing systems, and practical, a roadmap of technical and ethical challenges and future directions.
Load-bearing premise
The survey's comparative tables and performance conclusions rest on vendor-reported benchmark scores and repository documentation that the authors only partially verified with undocumented direct testing; if those reports are inaccurate, the comparisons lose their foundation.
Editorial extensions
If this is right
- If the taxonomy is adopted, systems can be compared on common dimensions rather than by vendor marketing claims, making capability gaps and trade-offs visible.
- The architectural pattern analysis gives system builders a decision tool: monolithic for reasoning coherence, pipeline for modularity, multi-agent for parallelization, hybrid for balance.
- The benchmark comparison establishes a baseline: commercial systems lead on HLE and GAIA, while open-source systems compete on cost, control, and domain-specific optimization.
- The roadmap of future directions (advanced reasoning, multimodal integration, domain specialization, human-AI collaboration, standardization) identifies what must happen for the field to mature.
Reading between the lines
- The same four-dimensional taxonomy could extend beyond Deep Research to any AI system that combines a model, tools, and a workflow, giving the broader agent ecosystem a shared vocabulary.
- The paper's mention of direct testing of public systems, without reporting details, suggests a verification protocol that future surveys could formalize and publish for reproducibility.
- The four architectural patterns and their failure-propagation characteristics could serve as a practical design heuristic even where benchmark data remain incomplete.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey proposes a hierarchical taxonomy of Deep Research systems organized along four technical dimensions (foundation models and reasoning engines, tool utilization and environmental interaction, task planning and execution control, and knowledge synthesis and output generation). It reviews more than 80 commercial and open-source systems, compares them across feature tables and benchmark scores (Tables 1-13), analyzes architectural patterns (monolithic, pipeline, multi-agent, hybrid), discusses implementation technologies, evaluation methodologies, applications, ethical considerations, and future directions, and provides a public repository of resources.
Significance. If its comparative data were reliable, this survey would be a useful reference contribution: it offers broad coverage of a fast-moving field, a clear organizing taxonomy, explicit system inclusion criteria in Section 5.5.1, and honest acknowledgment of absent benchmark evidence in several places (e.g., TREC in Section 5.1.2, FinEval in Section 5.3.2). The taxonomy is an external framing rather than an output of the surveyed systems, so circularity is not a concern. The paper's main weakness is that its analytical core—the cross-system benchmark comparison—is currently undercut by unresolved internal contradictions and by an undocumented verification procedure that is the only stated mechanism for resolving those contradictions.
major comments (4)
- [Section 3.3.1, Tables 8-9, Section 5.3.1] The manuscript contains directly contradictory numbers for the same systems and benchmarks. Table 8 reports Grok3Beta MMLU as 79.9%, while Table 9 reports 92.7% for the same system and source [299]; Section 5.3.1 states that OpenAI/DeepResearch averages 72.57% on GAIA, while Table 8 and Table 9 both report 67.36% pass@1 with source [197]. No explanation is given (e.g., different MMLU versions, different GAIA aggregation methods), and no correction is provided. Because these tables and passages constitute the paper's primary comparative evidence, the comparative claim is not currently supportable without reconciliation.
- [Section 5.5.3] The stated data-collection method 4, 'Experimental Verification: Where inconsistencies exist, we conducted direct testing of publicly available systems to verify capabilities,' is load-bearing precisely because the manuscript contains the inconsistencies identified above. However, no test dates, system versions, protocols, or results are reported anywhere. A reader therefore cannot determine which of the conflicting numbers is correct, and the paper's own method for resolving such conflicts is unverifiable as written.
- [Section 3.3.1, Table 8] The cross-system benchmark comparison is presented as evidence that commercial systems 'generally demonstrate leading performance,' but Table 8 has sparse and non-overlapping entries across systems: no single benchmark is populated for all or even most rows, several rows contain a single score, and the table mixes different benchmark families (HLE, MMLU, HotpotQA, GAIA) without a common evaluation context. The prose conclusion goes beyond what the table can support; claims should be restricted to pairwise or per-benchmark comparisons where data exist.
- [Section 5.5.4, Tables 8-12] The manuscript acknowledges in Section 5.5.4 that 'Systems undergo frequent updates, potentially rendering specific benchmark results obsolete,' yet none of the benchmark tables carry evaluation dates or version identifiers. Since the survey's snapshot is dated April 2025 and the ecosystem evolved rapidly during that period, the absence of dates makes it impossible to verify that the reported scores are mutually contemporaneous, which is a necessary condition for the comparisons in Tables 8-12 to be meaningful.
minor comments (5)
- [Section 3.3.1] There is a typo: 'HLE [212] hich measures' should be 'which measures'.
- [Section 1.4] The roadmap in Section 1.4 does not match the actual section numbering: implementation technologies are presented in Section 4, evaluation methodologies in Section 5, applications in Section 6, ethical considerations in Section 7, and future directions in Section 8, whereas the text refers to 'implementation technologies (Section 5), evaluation methodologies (Section 6), applications and use cases (Section 7), ethical considerations (Section 8), and future directions (Section 9)'.
- [Table 8] The footnote for GAIA ('GAIA Score(pass@1): Average score') is unclear; GAIA is typically reported as pass@1 averaged over the three difficulty levels, and the label should be aligned with the reporting convention used in the GAIA paper and in Table 9.
- [Table 12, Section 5.2.1] Response-time figures are inconsistent: Table 12 gives OpenAI/DeepResearch a range of 5-30 minutes while Section 5.2.1 says '5-10 minutes,' and Table 12 gives Perplexity/DeepResearch 2m59s while Section 5.2.1 says '2-5 minutes'; these should be reconciled or qualified by task complexity.
- [Table 8] Table 8 includes rows for Gemini-2.5 and Gemini-2.0-Flash, but Section 1.1's inclusion criteria target Deep Research systems, and the table does not clarify whether these rows refer to Gemini/DeepResearch or to the underlying models; this should be stated explicitly.
Circularity Check
No significant circularity: the survey's taxonomy is an external organizing framework, not a quantity derived from or fitted to the systems it surveys.
full rationale
This paper is a survey, not a predictive derivation: it proposes a four-dimensional taxonomy for describing Deep Research systems and then applies that taxonomy to organize architectural patterns, benchmarks, applications, and challenges. There is no equation, fitted parameter, or benchmark-derived quantity that is later relabeled as a prediction. The taxonomy is defined in Section 2 from technical capabilities, and the surveyed systems are then described in those terms; this is classification rather than derivation, so the central claim does not reduce to its inputs by construction. The paper does not rely on a load-bearing self-citation chain: no 'uniqueness theorem' or prior author result is invoked to make a choice forced, and the GitHub resource link is supporting material, not the foundation of the argument. The internal benchmark discrepancies noted by the skeptical reader (e.g., Grok3Beta MMLU 79.9% vs. 92.7%; GAIA 67.36% vs. 72.57%) are data-reliability or correctness concerns about vendor-reported metrics and undeclared direct testing, not circularity under any of the enumerated patterns. The paper itself flags data limitations and states that inconsistencies were checked, but even a failure to document those checks would affect empirical reliability, not circularity. Accordingly, an honest non-finding is appropriate: the survey is self-contained as an organizing framework, and no circular step can be exhibited from its text.
Assumptions & free parameters
free parameters (1)
- System inclusion criteria =
At least 2 of 3 core dimensions; public documentation; active development within past 12 months; representational…
assumptions (3)
- domain assumption Vendor-reported benchmarks and documentation accurately represent system capabilities
- ad hoc to paper The four proposed technical dimensions are the fundamental axes for categorizing Deep Research systems
- domain assumption Benchmark scores from different evaluations are comparable across systems despite differing protocols
Cite this review
Pith. "Pith review of A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications." pith.science (2026). https://pith.science/paper/YUO7CZMC
@misc{pith2026250612594,
author = {Pith},
title = {Pith review of: A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications},
year = {2026},
howpublished = {\url{https://pith.science/paper/YUO7CZMC}},
note = {Machine review of arXiv:2506.12594}
}
read the original abstract
This survey examines the rapidly evolving field of Deep Research systems -- AI-powered applications that automate complex research workflows through the integration of large language models, advanced information retrieval, and autonomous reasoning capabilities. We analyze more than 80 commercial and non-commercial implementations that have emerged since 2023, including OpenAI/Deep Research, Gemini/Deep Research, Perplexity/Deep Research, and numerous open-source alternatives. Through comprehensive examination, we propose a novel hierarchical taxonomy that categorizes systems according to four fundamental technical dimensions: foundation models and reasoning engines, tool utilization and environmental interaction, task planning and execution control, and knowledge synthesis and output generation. We explore the architectural patterns, implementation approaches, and domain-specific adaptations that characterize these systems across academic, scientific, business, and educational applications. Our analysis reveals both the significant capabilities of current implementations and the technical and ethical challenges they present regarding information accuracy, privacy, intellectual property, and accessibility. The survey concludes by identifying promising research directions in advanced reasoning architectures, multimodal integration, domain specialization, human-AI collaboration, and ecosystem standardization that will likely shape the future evolution of this transformative technology. By providing a comprehensive framework for understanding Deep Research systems, this survey contributes to both the theoretical understanding of AI-augmented knowledge work and the practical development of more capable, responsible, and accessible research technologies. The paper resources can be viewed at https://github.com/scienceaix/deepresearch.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 12 Pith papers
-
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
A 7B LLM agent trained with student-led distillation and one-step teacher corrections nearly matches a 72B teacher on reasoning and tool-use benchmarks.
-
HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research
A hierarchical evidence-graph benchmark reveals that multimodal deep-research models write fluent reports while failing citation, claim, and answer grounding.
-
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
An automatic pipeline evolves simple QA questions into 500 deep-research tasks with DAG-structured, fact-grounded rubrics that discriminate between models.
-
Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis
LLM deep-research agents rarely use historical analogies; a structural-decomposition plus cross-analogy-confirmation agent (CANA) sharply increases mechanism-grounded analogy claims and hidden-factor hits on the new A...
-
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
DEER uses 7 evaluation dimensions, 101 rubric items, task-specific expert guidance, and unsupported-claim backtracking to score deep-research reports; current systems score lowest on fulfilling expert requests and ana...
-
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
An automated red-teaming framework and 300-case benchmark show that a single unreliable website can induce unsafe responses in LLM search agents, with attack success rates up to 90.5%.
-
Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
In multilingual retrieval-augmented generation, models cite English evidence more accurately than translated evidence, and this language preference can outweigh document relevance.
-
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
A 20B autonomously reasoning deep-research agent trained with synthetic-data RL reaches 28.7% on Humanity's Last Exam, exceeding several larger and proprietary baselines.
-
Characterizing Deep Research: A Benchmark and Formal Definition
Deep research is characterized by high search and reasoning intensity; the new LiveDRBench measures claim-level precision and recall, where the best current model scores 0.55 F1.
-
FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality
A multi-LLM consensus pipeline turns 14,450 auto-generated candidate rubrics into 2,600 distinguishable gold rubrics that rank 10 financial deep-research systems from 58.58% to 22.23% pass rate.
-
DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent
A synthetic benchmark of 9,000 multi-hop web-research questions with difficulty tiers and teacher-generated search trajectories, plus an open-source RL training framework that reportedly lets 3B-parameter agents beat ...
-
Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
Branch-and-Browse, a tree-structured web agent with page-level action memory, reports 35.8% success and up to 40.4% less time than the cited Tree Search baseline on WebArena.
Reference graph
Works this paper leans on
-
[299]
Sarah Welsh. 2025. AI Benchmark Deep Dive: Gemini 2.5 and Humanity’s Last Exam. https://arize.com/blog/ai- benchmark-deep-dive-gemini-humanitys-last-exam/
2025
-
[197]
OpenAI. 2025. Introducing Deep Research. https://openai.com/index/introducing-deep-research/
2025
-
[1]
Adilzhan Adilkhanov, Amir Yelenov, Assylkhan Seitzhanov, Ayan Mazhitov, Azamat Abdikarimov, Danissa Sandyk- bayeva, Daryn Kenzhebek, Dinmukhammed Mukashev, Ilyas Umurbekov, Jabrail Chumakov, Kamila Spanova, Karina Burunchina, Madina Yergibay, Margulan Issa, Moldir Zabirova, Nurdaulet Zhuzbay, Nurlan Kabdyshev, Nurlan Zhani- yar, Rasul Yermagambet, Rustam ...
arXiv 2025
-
[2]
Agent-RL. 2024. ReSearch. https://github.com/Agent-RL/ReSearch
2024
-
[3]
Agno-AGI. 2025. Agno. https://github.com/agno-agi/agno
2025
-
[4]
Garima Agrawal, Sashank Gummuluri, and Cosimo Spera. 2024. Beyond-RAG: Question Identification and Answer Generation in Real-Time Conversations. arXiv:2410.10136 [cs.CL] https://arxiv.org/abs/2410.10136
arXiv 2024
-
[5]
Flowise AI. 2023. Flowise: Low-code LLM Application Building Tool. https://flowiseai.com/
2023
-
[6]
Nawaf Alampara, Mara Schilling-Wilhelmi, Martiño Ríos-García, Indrajeet Mandal, Pranav Khetarpal, Hargun Singh Grover, N. M. Anoop Krishnan, and Kevin Maik Jablonka. 2025. Probing the limitations of multimodal language models for chemistry and materials research. arXiv:2411.16955 [cs.LG] https://arxiv.org/abs/2411.16955
arXiv 2025
Show all 298 references
-
[7]
AlphaProof and AlphaGeometry teams. 2024. AI achieves silver-medal standard solving International Mathematical Olympiad problems. https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/
2024
-
[8]
Salaheddin Alzubi, Creston Brooks, Purva Chiniya, Edoardo Contente, Chiara von Gerlach, Lucas Irwin, Yihan Jiang, Arda Kaz, Windsor Nguyen, Sewoong Oh, Himanshu Tyagi, and Pramod Viswanath. 2025. Open Deep Search: Democratizing Search with Open-source Reasoning Agents. arXiv:2...
2025 arXiv
-
[9]
Lucio Anderlini, Matteo Barbetti, Giulio Bianchini, Diego Ciangottini, Stefano Dal Pra, Diego Michelotto, Carmelo Pellegrino, Rosa Petrini, Alessandro Pascolini, and Daniele Spiga. 2025. Supporting the development of Machine Learning for fundamental science in a federated Clou...
2025 arXiv
-
[10]
Mehrad Ansari and Seyed Mohamad Moosavi. 2023. Agent-based Learning of Materials Datasets from Scientific Literature. https://github.com/AI4ChemS/Eunomia. arXiv:2312.11690 [cs.AI] https://arxiv.org/abs/2312.11690
2023 arXiv
-
[11]
Anthropic. 2024. Building effective agents. https://www.anthropic.com/engineering/building-effective-agents
2024
-
[12]
Antropic. 2024. Model Context Protocol (MCP). https://docs.anthropic.com/en/docs/agents-and-tools/mcp
2024
-
[13]
Antropic. 2025. Claude takes research to new places. https://www.anthropic.com/news/research
2025
-
[14]
Prakash Aryan. 2024. LLMs as Debate Partners: Utilizing Genetic Algorithms and Adversarial Search for Adaptive Arguments. arXiv:2412.06229 [cs.AI] https://arxiv.org/abs/2412.06229
2024 arXiv
-
[15]
Johnson, Casey Dugan, and Michelle Bachman
Zahra Ashktorab, Qian Pan, Werner Geyer, Michael Desmond, Marina Danilevsky, James M. Johnson, Casey Dugan, and Michelle Bachman. 2024. Emerging Reliance Behaviors in Human-AI Text Generation: Hallucinations, Data Quality Assessment, and Cognitive Forcing Functions. arXiv:2409...
2024 arXiv
-
[16]
assafelovic. 2023. GPT-Researcher. https://github.com/assafelovic/gpt-researcher/
2023
-
[17]
Ahmet Yasin Aytar, Kemal Kilic, and Kamer Kaya. 2024. A Retrieval-Augmented Generation Framework for Academic Literature Navigation in Data Science. arXiv:2412.15404 [cs.IR] https://arxiv.org/abs/2412.15404
2024 arXiv
-
[18]
Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. 2025. ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. arXiv:2404.07738 [cs.CL] https://arxiv.org/ abs/2404.07738
2025 arXiv
-
[19]
Dzmitry Bahdanau, Nicolas Gontier, Gabriel Huang, Ehsan Kamalloo, Rafael Pardinas, Alex Piché, Torsten Scholak, Oleh Shliazhko, Jordan Prince Tremblay, Karam Ghanem, Soham Parikh, Mitul Tiwari, and Quaizar Vohra. 2024. TapeAgents: a Holistic Framework for Agent Development and...
2024 arXiv
-
[20]
Gal Bakal, Ali Dasdan, Yaniv Katz, Michael Kaufman, and Guy Levin. 2025. Experience with GitHub Copilot for Developer Productivity at Zoominfo. arXiv:2501.13282 [cs.SE] https://arxiv.org/abs/2501.13282
2025 arXiv
-
[21]
Howard Balshem, Mark Helfand, Holger J Schünemann, Andrew D Oxman, Regina Kunzand Jan Brozek, Gunn E Vist, Yngve Falck-Ytter, Joerg Meerpohl, Susan Norris, and Gordon H Guyatt. 2011. GRADE guidelines: 3. Rating the quality of evidence. https://pubmed.ncbi.nlm.nih.gov/21208779/
2011
-
[22]
Samuel Barham, Orion Weller, Michelle Yuan, Kenton Murray, Mahsa Yarmohammadi, Zhengping Jiang, Siddharth Vashishtha, Alexander Martin, Anqi Liu, Aaron Steven White, Jordan Boyd-Graber, and Benjamin Van Durme
-
[23]
Rhea Basappa, Mustafa Tekman, Hong Lu, Benjamin Faught, Sandeep Kakar, and Ashok K. Goel. 2024.Social AI Agents Too Need to Explain Themselves. Springer Nature Switzerland, 351–360. doi:10.1007/978-3-031-63028-6_29
2024 doi
-
[24]
Joeran Beel, Min-Yen Kan, and Moritz Baumgart. 2025. Evaluating Sakana’s AI Scientist for Autonomous Research: Wishful Thinking or an Emerging Reality Towards ’Artificial Research Intelligence’ (ARI)? arXiv:2502.14297 [cs.IR] https://arxiv.org/abs/2502.14297
2025
-
[25]
Morad Behandish, John Maxwell III, and Johan de Kleer. 2022. AI Research Associate for Early-Stage Scientific Discovery. arXiv:2202.03199 [cs.AI] https://arxiv.org/abs/2202.03199
2022 arXiv
-
[26]
Yoshua Bengio, Michael Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, Sören Mindermann, Adam Oberman, Jesse Richardson, Oliver Richardson, Marc-Antoine Rondeau, Pierre-Luc St-Charles, and David Williams-King. 2025. Superintelligent Agents Pose Catas...
2025 arXiv
-
[27]
Karim Benharrak, Tim Zindulka, and Daniel Buschek. 2024. Deceptive Patterns of Intelligent and Interactive Writing Assistants. arXiv:2404.09375 [cs.HC] https://arxiv.org/abs/2404.09375
2024 arXiv
-
[28]
Zhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, and Huajun Chen. 2024. OceanGPT: A Large Language Model for Ocean Science Tasks. http://oceangpt.zjukg.cn/. arXiv:2310.02031 [cs.CL] https: //arxiv.org/abs/2310.02031
2024 arXiv
-
[29]
Stefano Bianchini, Moritz Müller, and Pierre Pelletier. 2024. Drivers and Barriers of AI Adoption and Use in Scientific Research. arXiv:2312.09843 [cs.CY] https://arxiv.org/abs/2312.09843
2024 arXiv
-
[30]
bindAI. 2025. ChatGPT Deep Research vs Perplexity – Which One Is Better? https://blog.getbind.co/2025/02/03/ chatgpt-deep-research-is-it-better-than-perplexity/
2025
-
[31]
Francisco Bolanos, Angelo Salatino, Francesco Osborne, and Enrico Motta. 2024. Artificial Intelligence for Literature Reviews: Opportunities and Challenges. arXiv:2402.08565 [cs.AI] https://arxiv.org/abs/2402.08565
2024 arXiv
-
[32]
Bolt. 2024. Bolt. https://bolt.new/
2024
-
[33]
bracai. 2025. MMLU benchmark: Testing LLMs multi-task capabilities. https://www.bracai.eu/post/mmlu-benchmark
2025
-
[34]
Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. 2023. ChemCrow: Augmenting large-language models with chemistry tools. arXiv:2304.05376 [physics.chem-ph] https://arxiv.org/abs/ 2304.05376
2023 arXiv
-
[35]
Chris Brown and Jason Cusati. 2024. Exploring the Evidence-Based Beliefs and Behaviors of LLM-Based Programming Assistants. arXiv:2407.13900 [cs.SE] https://arxiv.org/abs/2407.13900
2024 arXiv
-
[36]
browserbase. 2025. Open-operator. https://github.com/browserbase/open-operator
2025
-
[37]
btahir. 2024. open_deep_research. https://github.com/btahir/open-deep-research
2024
-
[38]
ByteDance. 2024. Coze Space. https://www.coze.cn/space-preview
2024
-
[39]
ByteDance. 2025. agent-tars. https://github.com/bytedance/UI-TARS-desktop/tree/main/apps/agent-tars
2025
-
[40]
Beatriz Cabrero-Daniel, Tomas Herda, Victoria Pichler, and Martin Eder. 2024. Exploring Human-AI Collaboration in Agile: Customised LLM Meeting Assistants. arXiv:2404.14871 [cs.SE] https://arxiv.org/abs/2404.14871
2024 arXiv
-
[41]
Filipe Calegario, Vanilson Burégio, Francisco Erivaldo, Daniel Moraes Costa Andrade, Kailane Felix, Nathalia Barbosa, Pedro Lucas da Silva Lucena, and César França. 2023. Exploring the intersection of Generative AI and Software Development. arXiv:2312.14262 [cs.SE] https://arx...
2023 arXiv
-
[42]
Nicholas Camara. 2025. open-deep-research. https://github.com/nickscamara/open-deep-research
2025
-
[43]
Camel AI. 2025. OWL. https://github.com/camel-ai/owl
2025
-
[44]
Mustafa Rafique, Eliu Huerta, Bo Li, Ian Foster, and Rick Stevens
Franck Cappello, Sandeep Madireddy, Robert Underwood, Neil Getty, Nicholas Lee-Ping Chia, Nesar Ramachandra, Josh Nguyen, Murat Keceli, Tanwi Mallick, Zilinghan Li, Marieme Ngom, Chenhui Zhang, Angel Yanguas-Gil, Evan Antoniuk, Bhavya Kailkhura, Minyang Tian, Yufeng Du, Yuan-S...
2025 arXiv
-
[45]
Peter Cardon, Carolin Fleischmann, Jolanta Aritz, Minna Logemann, and Jeanette Heidewald. 2023. The Challenges and Opportunities of AI-Assisted Writing: Developing AI Literacy for the AI Age. https://journals.sagepub.com/doi/ abs/10.1177/23294906231176517
2023 doi
-
[46]
Pan, Shuyi Yang, Lakshya A
Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica. 2025. Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657 [cs.AI] h...
2025 arXiv
-
[47]
Eric Chamoun, Michael Schlichktrull, and Andreas Vlachos. 2024. Automated Focused Feedback Generation for Scientific Writing Assistance. arXiv:2405.20477 [cs.CL] https://arxiv.org/abs/2405.20477
2024 arXiv
-
[48]
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate. https://github.com/thunlp/ChatEval. arXiv:2308.07201 [cs.CL] https://arxiv.org/abs/2308.07201 A...
2023 arXiv
-
[49]
Jiangjie Chen, Xintao Wang, Rui Xu, Siyu Yuan, Yikai Zhang, Wei Shi, Jian Xie, Shuang Li, Ruihan Yang, Tinghui Zhu, Aili Chen, Nianqi Li, Lida Chen, Caiyu Hu, Siye Wu, Scott Ren, Ziquan Fu, and Yanghua Xiao. 2024. From Persona to Personalization: A Survey on Role-Playing Langu...
2024 arXiv
-
[50]
Kexin Chen, Hanqun Cao, Junyou Li, Yuyang Du, Menghao Guo, Xin Zeng, Lanqing Li, Jiezhong Qiu, Pheng Ann Heng, and Guangyong Chen. 2024. An Autonomous Large Language Model Agent for Chemical Literature Data Mining. arXiv:2402.12993 [cs.IR] https://arxiv.org/abs/2402.12993
2024 arXiv
-
[51]
Pengcheng Chen, Jin Ye, Guoan Wang, Yanjun Li, Zhongying Deng, Wei Li, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, Benyou Wang, Shaoting Zhang, Bin Fu, Jianfei Cai, Bohan Zhuang, Eric J Seibel, Junjun He, and Yu Qiao
-
[52]
Tingting Chen, Srinivas Anumasa, Beibei Lin, Vedant Shah, Anirudh Goyal, and Dianbo Liu. 2025. Auto- Bench: An Automated Benchmark for Scientific Discovery in LLMs. https://github.com/AutoBench/AutoBench. arXiv:2502.15224 [cs.LG] https://arxiv.org/abs/2502.15224
2025 arXiv
-
[53]
Valerie Chen, Alan Zhu, Sebastian Zhao, Hussein Mozannar, David Sontag, and Ameet Talwalkar. 2025. Need Help? Designing Proactive AI Assistants for Programming. arXiv:2410.04596 [cs.HC] https://arxiv.org/abs/2410.04596
2025 arXiv
-
[54]
Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. 2023. AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent B...
2023 arXiv
-
[55]
Qinyuan Cheng, Tianxiang Sun, Xiangyang Liu, Wenwei Zhang, Zhangyue Yin, Shimin Li, Linyang Li, Zhengfu He, Kai Chen, and Xipeng Qiu. 2024. Can AI Assistants Know What They Don’t Know? arXiv:2401.13275 [cs.CL] https://arxiv.org/abs/2401.13275
2024 arXiv
-
[56]
ZhaoCheng, DianeWan, MatthewAbueg,Sahra Ghalebikesabi, RenYi, EugeneBagdasarian,Borja Balle,Stefan Mellem, and Shawn O’Banion. 2024. CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data. https: //www.aimodels.fyi/papers/arxiv/ci-bench-benchmarking-con...
2024 arXiv
-
[57]
Hen- ley
Bhavya Chopra, Ananya Singha, Anna Fariha, Sumit Gulwani, Chris Parnin, Ashish Tiwari, and Austin Z. Hen- ley. 2023. Conversational Challenges in AI-Powered Data Science: Obstacles, Needs, and Design Opportunities. arXiv:2310.16164 [cs.HC] https://arxiv.org/abs/2310.16164
2023 arXiv
-
[58]
Daniel J. H. Chung, Zhiqi Gao, Yurii Kvasiuk, Tianyi Li, Moritz Münchmeyer, Maja Rudolph, Frederic Sala, and Sai Chaitanya Tadepalli. 2025. Theoretical Physics Benchmark (TPBench) – a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics. https://tpbench.org/. ...
2025 arXiv
-
[59]
Umut Cihan, Vahid Haratian, Arda İçöz, Mert Kaan Gül, Ömercan Devran, Emircan Furkan Bayendur, Baykal Mehmet Uçar, and Eray Tüzün. 2024. Automated Code Review In Practice. arXiv:2412.18531 [cs.SE] https://arxiv.org/abs/ 2412.18531
2024 arXiv
-
[60]
Dave Citron. 2025. Deep Research is now available on Gemini 2.5 Pro Experimental. https://blog.google/products/ gemini/deep-research-gemini-2-5-pro-experimental/
2025
-
[61]
Cline. 2024. Cline. https://github.com/cline/cline
2024
-
[62]
2025.Devin.ai
Cognition Labs. 2025.Devin.ai. https://devin.ai
2025
-
[63]
Consensus. 2025. Consensus. https://consensus.app/
2025
-
[64]
crewAIInc. 2023. CrewAI. https://github.com/crewAIInc/crewAI
2023
-
[65]
Cursor. 2023. Cursor. https://www.cursor.com/
2023
-
[66]
Danesh, Tu Trinh, Benjamin Plaut, and Nguyen X
Mohamad H. Danesh, Tu Trinh, Benjamin Plaut, and Nguyen X. Khanh. 2025. Learning to Coordinate with Experts. https://github.com/modanesh/YRC-Bench. arXiv:2502.09583 [cs.LG] https://arxiv.org/abs/2502.09583
2025
-
[67]
Kristin M. de Payrebrune, Kathrin Flaßkamp, Tom Ströhla, Thomas Sattel, Dieter Bestle, Benedict Röder, Peter Eberhard, Sebastian Peitz, Marcus Stoffel, Gulakala Rutwik, Borse Aditya, Meike Wohlleben, Walter Sextro, Maximilian Raff, C. David Remy, Manish Yadav, Merten Stender, ...
2024 arXiv
-
[68]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, et al. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12...
2025 arXiv
-
[69]
Akash Dhruv and Anshu Dubey. 2025. Leveraging Large Language Models for Code Translation and Software Development in Scientific Computing. arXiv:2410.24119 [cs.SE] https://arxiv.org/abs/2410.24119 84 Xu et al
2025
-
[70]
Talissa Dreossi. 2025. Bridging Logic Programming and Deep Learning for Explainability through ILASP.Electronic Proceedings in Theoretical Computer Science416 (Feb. 2025), 314–323. doi:10.4204/eptcs.416.31
2025 doi
-
[71]
It makes you think
Ian Drosos, Advait Sarkar, Xiaotong Xu, and Neil Toronto. 2025. "It makes you think": Provocations Help Restore Critical Thinking to AI-Assisted Knowledge Work. arXiv:2501.17247 [cs.HC] https://arxiv.org/abs/2501.17247
2025 arXiv
-
[72]
Omer Dunay, Daniel Cheng, Adam Tait, Parth Thakkar, Peter C Rigby, Andy Chiu, Imad Ahmad, Arun Ganesan, Chandra Maddila, Vijayaraghavan Murali, Ali Tayyebi, and Nachiappan Nagappan. 2024. Multi-line AI-assisted Code Authoring. arXiv:2402.04141 [cs.SE] https://arxiv.org/abs/2402.04141
2024 arXiv
-
[73]
Steffen Eger, Yong Cao, Jennifer D’Souza, Andreas Geiger, Christian Greisinger, Stephanie Gross, Yufang Hou, Brigitte Krenn, Anne Lauscher, Yizhi Li, Chenghua Lin, Nafise Sadat Moosavi, Wei Zhao, and Tristan Miller. 2025. Transforming Science with Large Language Models: A Surv...
2025
-
[74]
Elicit. 2025. Elicit. https://elicit.com/?redirected=true
2025
-
[75]
Michael D. Ernst. 2017. Natural Language is a Programming Language: Applying Natural Language Processing to Software Development. https://drops.dagstuhl.de/storage/00lipics/lipics-vol071-snapl2017/LIPIcs.SNAPL.2017.4/ LIPIcs.SNAPL.2017.4.pdf
2017
-
[76]
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models. arXiv:2405.06211 [cs.CL] https://arxiv.org/abs/2405.06211
2024 arXiv
-
[77]
Flowith. 2024. Flowith Oracle Mode. https://flowith.net/
2024
-
[78]
Forethought-Technologies. 2023. AutoChain. https://github.com/Forethought-Technologies/AutoChain
2023
-
[79]
César França. 2023. AI empowering research: 10 ways how science can benefit from AI. arXiv:2307.10265 [cs.GL] https://arxiv.org/abs/2307.10265
2023 arXiv
-
[80]
Future-House. 2023. PaperQA. https://github.com/Future-House/paper-qa
2023
-
[81]
GAIR-NLP. 2025. DeepResearcher. https://github.com/GAIR-NLP/DeepResearcher
2025
-
[82]
Difei Gao, Lei Ji, Luowei Zhou, Kevin Qinghong Lin, Joya Chen, Zihan Fan, and Mike Zheng Shou. 2023. AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn. arXiv:2306.08640 [cs.CV] https: //arxiv.org/abs/2306.08640
2023 arXiv
-
[83]
Shanghua Gao, Ada Fang, Yepeng Huang, Valentina Giunchiglia, Ayush Noori, Jonathan Richard Schwarz, Yasha Ektefaie, Jovana Kondic, and Marinka Zitnik. 2024. Empowering Biomedical Discovery with AI Agents. arXiv:2404.02831 [cs.AI] https://arxiv.org/abs/2404.02831
2024 arXiv
-
[84]
Alireza Ghafarollahi and Markus J. Buehler. 2024. SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning. arXiv:2409.05556 [cs.AI] https://arxiv.org/abs/2409.05556
2024 arXiv
-
[85]
Luca Gioacchini, Marco Mellia, Idilio Drago, Alexander Delsanto, Giuseppe Siracusano, and Roberto Bifulco. 2024. AutoPenBench: Benchmarking Generative Agents for Penetration Testing. https://github.com/lucagioacchini/auto- pen-bench. arXiv:2410.03225 [cs.CR] https://arxiv.org/...
2024 arXiv
-
[86]
Github. 2021. Github Copilot. https://github.com/features/copilot?ref=nav.poetries.top
2021
-
[87]
Amr Gomaa, Michael Sargious, and Antonio Krüger. 2024. AdaptoML-UX: An Adaptive User-centered GUI-based AutoML Toolkit for Non-AI Experts and HCI Researchers. https://github.com/MichaelSargious/AdaptoML_UX. arXiv:2410.17469 [cs.HC] https://arxiv.org/abs/2410.17469
2024 arXiv
-
[88]
Google. 2021. BIG-bench. https://github.com/google/BIG-bench
2021
-
[89]
Google. 2024. Try Deep Research and our new experimental model in Gemini, your AI assistant. https://blog.google/ products/gemini/google-gemini-deep-research/
2024
-
[90]
Google. 2025. A2A. https://github.com/google/A2A
2025
-
[91]
Google. 2025. Agent Development Kit. https://google.github.io/adk-docs/
2025
-
[92]
Google. 2025. Announcing the Agent2Agent Protocol (A2A). https://developers.googleblog.com/en/a2a-a-new-era-of- agent-interoperability/
2025
-
[93]
Google. 2025. Gemini 2.0 Flash (Feb ’25): Intelligence, Performance and Price Analysis. https://artificialanalysis.ai/ models/gemini-2-0-flash
2025
-
[94]
Google. 2025. gemini-fullstack-langgraph-quickstart. https://github.com/google-gemini/gemini-fullstack-langgraph- quickstart
2025
-
[95]
Google. 2025. NotebookLm. https://notebooklm.google/
2025
-
[96]
Kanika Goswami, Puneet Mathur, Ryan Rossi, and Franck Dernoncourt. 2025. ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution. arXiv:2502.00989 [cs.CL] https://arxiv.org/abs/2502.00989
2025 arXiv
-
[97]
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Ko...
2025
-
[98]
Gower, Konstantin Korovin, Daniel Brunnsåker, Filip Kronström, Gabriel K
Alexander H. Gower, Konstantin Korovin, Daniel Brunnsåker, Filip Kronström, Gabriel K. Reder, Ievgeniia A. Tiukova, Ronald S. Reiserer, John P. Wikswo, and Ross D. King. 2024. The Use of AI-Robotic Systems for Scientific Discovery. arXiv:2406.17835 [cs.LG] https://arxiv.org/ab...
2024
-
[99]
Tianyang Gu, Jingjin Wang, Zhihao Zhang, and HaoHong Li. 2025. LLMs can Realize Combinatorial Creativity: Generating Creative Ideas via LLMs for Scientific Research. arXiv:2412.14141 [cs.AI] https://arxiv.org/abs/2412.14141
2025 arXiv
-
[100]
Yuzhe Gu, Wenwei Zhang, Chengqi Lyu, Dahua Lin, and Kai Chen. 2025. Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs. arXiv:2503.02846 [cs.CL] https://arxiv.org/abs/2503.02846
2025 arXiv
-
[101]
Chengquan Guo, Xun Liu, Chulin Xie, Andy Zhou, Yi Zeng, Zinan Lin, Dawn Song, and Bo Li. 2024. Red- Code: Risky Code Execution and Generation Benchmark for Code Agents. https://github.com/AI-secure/RedCode. arXiv:2411.07781 [cs.SE] https://arxiv.org/abs/2411.07781
2024 arXiv
-
[102]
Siyuan Guo, Cheng Deng, Ying Wen, Hechang Chen, Yi Chang, and Jun Wang. 2024. DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning. https://github.com/guosyjlu/DS-Agent. arXiv:2402.17453 [cs.LG] https://arxiv.org/abs/2402.17453
2024 arXiv
-
[103]
Xin Guo, Haotian Xia, Zhaowei Liu, Hanyang Cao, Zhi Yang, Zhiqiang Liu, Sizhe Wang, Jinyi Niu, Chuqi Wang, Yanhui Wang, Xiaolong Liang, Xiaoming Huang, Bing Zhu, Zhongyu Wei, Yun Chen, Weining Shen, and Liwen Zhang. 2024. FinEval: A Chinese Financial Domain Knowledge Evaluatio...
2024 arXiv
-
[104]
Hilda Hadan, Derrick Wang, Reza Hadi Mogavi, Joseph Tu, Leah Zhang-Kennedy, and Lennart E. Nacke. 2024. The Great AI Witch Hunt: Reviewers Perception and (Mis)Conception of Generative AI in Research Writing. https: //arxiv.org/abs/2407.12015
2024
-
[105]
Sukjin Han. 2024. Mining Causality: AI-Assisted Search for Instrumental Variables. arXiv:2409.14202 [econ.EM] https://arxiv.org/abs/2409.14202
2024 arXiv
-
[106]
LaToza, and Brittany Johnson
Ebtesam Al Haque, Chris Brown, Thomas D. LaToza, and Brittany Johnson. 2025. Towards Decoding Developer Cognition in the Age of AI Assistants. arXiv:2501.02684 [cs.HC] https://arxiv.org/abs/2501.02684
2025 arXiv
-
[107]
Gaole He, Patrick Hemmer, Michael Vössing, Max Schemmer, and Ujwal Gadiraju. 2025. Fine-Grained Appropriate Reliance: Human-AI Collaboration with a Multi-Step Transparent Decision Workflow for Complex Task Decomposition. arXiv:2501.10909 [cs.AI] https://arxiv.org/abs/2501.10909
2025 arXiv
-
[108]
Kaveen Hiniduma, Suren Byna, Jean Luca Bez, and Ravi Madduri. 2024. AI Data Readiness Inspector (AIDRIN) for Quantitative Assessment of Data Readiness for AI. InProceedings of the 36th International Conference on Scientific and Statistical Database Management (SSDBM 2024). ACM...
2024
-
[109]
HKUDS. 2025. AI-Researcher. https://github.com/HKUDS/AI-Researcher
2025
-
[110]
Brendan Hogan, Anmol Kabra, Felipe Siqueira Pacheco, Laura Greenstreet, Joshua Fan, Aaron Ferber, Marta Ummus, Alecsander Brito, Olivia Graham, Lillian Aoki, Drew Harvell, Alex Flecker, and Carla Gomes. 2024. AiSciVision: A Framework for Specializing Large Multimodal Models in...
2024 arXiv
-
[111]
Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024. MetaGPT: Meta Programming for A Multi-Agent Collaborative...
2024 arXiv
-
[112]
Hong Kong University Data Science Lab. 2024. Auto-Deep-Research. https://github.com/HKUDS/Auto-Deep- Research
2024
-
[113]
Betty Li Hou, Kejian Shi, Jason Phang, James Aung, Steven Adler, and Rosie Campbell. 2024. Large Language Models as Misleading Assistants in Conversation. arXiv:2407.11789 [cs.CL] https://arxiv.org/abs/2407.11789
2024 arXiv
-
[114]
Shulin Huang, Shirong Ma, Yinghui Li, Mengzuo Huang, Wuhe Zou, Weidong Zhang, and Hai-Tao Zheng. 2024. LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles. https://github.com/THUKElab/LatEval. arXiv:2308.10855 [cs.CL] htt...
2024 arXiv
-
[115]
HuggingFace. 2025. smolagents: open_deep_research. https://github.com/huggingface/smolagents/tree/main/ examples/open_deep_research
2025
-
[116]
Faria Huq, Abdus Samee, David Chuan-En Lin, Alice Xiaodi Tang, and Jeffrey P Bigham. 2025. NoTeeline: Supporting Real-Time, Personalized Notetaking with LLM-Enhanced Micronotes. InProceedings of the 30th International Conference on Intelligent User Interfaces (IUI ’25). ACM, 1...
2025
-
[117]
Kurando IIDA and Kenjiro MIMURA. 2024. CATER: Leveraging LLM to Pioneer a Multidimensional, Reference- Independent Paradigm in Translation Quality Evaluation. arXiv:2412.11261 [cs.CL] https://arxiv.org/abs/2412.11261 86 Xu et al
2024 arXiv
-
[118]
Seyed Mohammad Ali Jafari. 2024. Streamlining the Selection Phase of Systematic Literature Reviews (SLRs) Using AI-Enabled GPT-4 Assistant API. arXiv:2402.18582 [cs.DL] https://arxiv.org/abs/2402.18582
2024 arXiv
-
[119]
Rishab Jain and Aditya Jain. 2023. Generative AI in Writing Research Papers: A New Type of Algorithmic Bias and Uncertainty in Scholarly Work. arXiv:2312.10057 [cs.CY] https://arxiv.org/abs/2312.10057
2023 arXiv
-
[120]
Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, and Yuxiang Wu. 2025. AIDE: AI-Driven Exploration in the Space of Code. arXiv:2502.13138 [cs.AI] https://arxiv.org/abs/2502.13138
2025 arXiv
-
[121]
Jina AI. 2025. node-DeepResearch. https://github.com/jina-ai/node-DeepResearch
2025
-
[122]
Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu. 2025. DSBench: How Far Are Data Science Agents from Becoming Data Science Experts? https: //github.com/LiqiangJing/DSBench. arXiv:2409.07703 [cs.AI] https://arxi...
2025 arXiv
-
[123]
Nicola Jones. 2025. OpenAI’s ‘deep research’ tool: is it useful for scientists? https://www.nature.com/articles/d41586- 025-00377-9
2025
-
[124]
Vijay Joshi and Iver Band. 2024. Disrupting Test Development with AI Assistants: Building the Base of the Test Pyramid with Three AI Coding Assistants. (Oct. 2024). doi:10.36227/techrxiv.173014488.82191966/v1
2024
-
[125]
Majeed Kazemitabaar, Jack Williams, Ian Drosos, Tovi Grossman, Austin Zachary Henley, Carina Negreanu, and Advait Sarkar. 2024. Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition. In Proceedings of the 37th Annual ACM Symposium...
2024
-
[126]
CTOL Editors Ken. 2025. Gemini Launches Deep Research on 2.5 Pro Aiming to Redefine AI-Powered Analysis with Strong Lead Over OpenAI. https://www.ctol.digital/news/gemini-deep-research-launch-2-5-pro-vs-openai/
2025
-
[127]
Antti Keurulainen, Isak Westerlund, Samuel Kaski, and Alexander Ilin. 2021. Learning to Assist Agents by Observing Them. arXiv:2110.01311 [cs.AI] https://arxiv.org/abs/2110.01311
2021 arXiv
-
[128]
Abdullah Khalili and Abdelhamid Bouchachia. 2022. Toward Building Science Discovery Machines. arXiv:2103.15551 [cs.AI] https://arxiv.org/abs/2103.15551
2022 arXiv
-
[129]
Stefan Kramer, Mattia Cerrato, Sašo Džeroski, and Ross King. 2023. Automated Scientific Discovery: From Equation Discovery to Autonomous Discovery Systems. arXiv:2305.02251 [cs.AI] https://arxiv.org/abs/2305.02251
2023 arXiv
-
[130]
Kummerfeld, Anne Lauscher, Kevin Leyton-Brown, Sheng Lu, Mausam, Margot Mieskes, Aurélie Névéol, Danish Pruthi, Lizhen Qu, Roy Schwartz, Noah A
Ilia Kuznetsov, Osama Mohammed Afzal, Koen Dercksen, Nils Dycke, Alexander Goldberg, Tom Hope, Dirk Hovy, Jonathan K. Kummerfeld, Anne Lauscher, Kevin Leyton-Brown, Sheng Lu, Mausam, Margot Mieskes, Aurélie Névéol, Danish Pruthi, Lizhen Qu, Roy Schwartz, Noah A. Smith, Thamar ...
2024 arXiv
-
[131]
Martin Lance. 2024. open_deep_research. https://github.com/langchain-ai/open_deep_research
2024
-
[132]
Hao Lang, Fei Huang, and Yongbin Li. 2025. Debate Helps Weak-to-Strong Generalization. arXiv:2501.13124 [cs.CL] https://arxiv.org/abs/2501.13124
2025 arXiv
-
[133]
LangChain. 2025. How to think about agent frameworks. https://blog.langchain.dev/how-to-think-about-agent- frameworks/. https://docs.google.com/spreadsheets/d/1B37VxTBuGLeTSPVWtz7UMsCdtXrqV5hCjWkbHN8tfAo/
2025
-
[134]
langChain AI. 2024. LangGraph. https://github.com/langchain-ai/langgraph
2024
-
[135]
Andrew Laverick, Kristen Surrao, Inigo Zubeldia, Boris Bolliet, Miles Cranmer, Antony Lewis, Blake Sherwin, and Julien Lesgourgues. 2024. Multi-Agent System for Cosmological Parameter Analysis. arXiv:2412.00431 [astro-ph.IM] https://arxiv.org/abs/2412.00431
2024 arXiv
-
[136]
Eunhae Lee. 2024. Towards Ethical Personal AI Applications: Practical Considerations for AI Assistants with Long-Term Memory. arXiv:2409.11192 [cs.CY] https://arxiv.org/abs/2409.11192
2024 arXiv
-
[137]
Yuho Lee, Taewon Yun, Jason Cai, Hang Su, and Hwanjun Song. 2024. UniSumEval: Towards Unified, Fine- Grained, Multi-Dimensional Summarization Evaluation for LLMs. https://github.com/DISL-Lab/UniSumEval-v1.0. arXiv:2409.19898 [cs.CL] https://arxiv.org/abs/2409.19898
2024 arXiv
-
[138]
Letta-AI. 2023. Letta. https://github.com/letta-ai/letta
2023
-
[139]
Berger, and Stephen N
Kyla Levin, Nicolas van Kempen, Emery D. Berger, and Stephen N. Freund. 2025. ChatDBG: An AI-Powered Debugging Assistant. arXiv:2403.16354 [cs.SE] https://arxiv.org/abs/2403.16354
2025 arXiv
-
[140]
James R. Lewis. 2018. The System Usability Scale: Past, Present, and Future. International Journal of Human–Computer Interaction 34, 7 (2018), 577–590. doi:10.1080/10447318.2018.1455307 arXiv:https://doi.org/10.1080/10447318.2018.1455307
2018
-
[141]
Li, Been Kim, and Zi Wang
Belinda Z. Li, Been Kim, and Zi Wang. 2025. QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks? arXiv:2503.22674 [cs.AI] https://arxiv.org/abs/2503.22674
2025
-
[142]
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. arXiv:2303.17760 [cs.AI] https: //arxiv.org/abs/2303.17760 A Comprehensive Survey of Deep Resea...
2023 arXiv
-
[143]
Jiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube, Bingsheng Yao, Xuhai Xu, Dakuo Wang, Elizabeth Mynatt, and Varun Mishra. 2025. Vital Insight: Assisting Experts’ Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LL...
2025
-
[144]
Xuefeng Li, Haoyang Zou, and Pengfei Liu. 2025. TORL: Scaling Tool-Integrated RL. https://github.com/GAIR- NLP/ToRL. https://arxiv.org/pdf/2503.23383
2025 arXiv
-
[145]
Yuan Li, Yixuan Zhang, and Lichao Sun. 2023. MetaAgents: Simulating Interactions of Human Behaviors for LLM-based Task-oriented Coordination via Collaborative Generative Agents. arXiv:2310.06500 [cs.AI] https://arxiv.org/abs/2310. 06500
2023 arXiv
-
[146]
Zhuoyan Li, Chen Liang, Jing Peng, and Ming Yin. 2024. How Does the Disclosure of AI Assistance Affect the Perceptions of Writing? arXiv:2410.04545 [cs.CL] https://arxiv.org/abs/2410.04545
2024 arXiv
-
[147]
Wilson, Woosang Lim, and William Yang Wang
Zekun Li, Xianjun Yang, Kyuri Choi, Wanrong Zhu, Ryan Hsieh, HyeonJung Kim, Jin Hyuk Lim, Sungyoung Ji, Byungju Lee, Xifeng Yan, Linda Ruth Petzold, Stephen D. Wilson, Woosang Lim, and William Yang Wang. 2025. MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Sci...
2025 arXiv
-
[148]
Liang, Chenyang Yang, and Brad A
Jenny T. Liang, Chenyang Yang, and Brad A. Myers. 2023. A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and Challenges. arXiv:2303.17125 [cs.SE] https://arxiv.org/abs/2303.17125
2023 arXiv
-
[149]
Manning, Christopher Ré, Diana Acosta-Navas, Drew A
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...
2023 arXiv
-
[150]
Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and Xiaodong Shi. 2023. Automated scholarly paper review: Concepts, technologies, and challenges.Information Fusion98 (Oct. 2023), 101830. doi:10.1016/j.inffus.2023.101830
2023
-
[151]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv:2109.07958 [cs.CL] https://arxiv.org/abs/2109.07958
2022 arXiv
-
[152]
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, Jianfeng Gao, and Chunyuan Li. 2023. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents. arXiv:2311.05437 [cs.CV] https://arxiv.org/abs/2311.05437
2023 arXiv
-
[154]
Xiao Liu, Bo Qin, Dongzhu Liang, Guang Dong, Hanyu Lai, Hanchen Zhang, Hanlin Zhao, Iat Long Iong, Jiadai Sun, Jiaqi Wang, Junjie Gao, Junjun Shan, Kangning Liu, Shudan Zhang, Shuntian Yao, Siyi Cheng, Wentao Yao, Wenyi Zhao, Xinghan Liu, Xinyi Liu, Xinying Chen, Xinyue Yang, ...
2024 arXiv
-
[155]
Zijun Liu, Kaiming Liu, Yiqi Zhu, Xuanyu Lei, Zonghan Yang, Zhenhe Zhang, Peng Li, and Yang Liu. 2024. AIGS: Generating Science from AI-Powered Automated Falsification. arXiv:2411.11910 [cs.LG] https://arxiv.org/abs/2411. 11910
2024 arXiv
-
[156]
Zhiwei Liu, Weiran Yao, Jianguo Zhang, Le Xue, Shelby Heinecke, Rithesh Murthy, Yihao Feng, Zeyuan Chen, Juan Carlos Niebles, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese. 2023. BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Ag...
2023 arXiv
-
[157]
Renze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Jian Xie, Yuxuan Sun, Yusen Zhang, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Wenchao Ma, Xi Li, Kai Zhang, Congying Xia, Lifu Huang, and Wenpeng Yin. 2025. AAAR-1.0: Assessing AI’s Potential to Assist R...
2025 arXiv
-
[158]
Cong Lu, Shengran Hu, and Jeff Clune. 2025. Automated Capability Discovery via Model Self-Exploration. arXiv:2502.07577 [cs.LG] https://arxiv.org/abs/2502.07577 88 Xu et al
2025 arXiv
-
[159]
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292 [cs.AI] https://arxiv.org/abs/2408.06292
2024 arXiv
-
[160]
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. https://scienceqa.github.io/. arXiv:2209.09513 [cs.CL] htt...
2022 arXiv
-
[161]
Chandra Maddila, Negar Ghorbani, Kosay Jabre, Vijayaraghavan Murali, Edwin Kim, Parth Thakkar, Nikolay Pavlovich Laptev, Olivia Harman, Diana Hsu, Rui Abreu, and Peter C. Rigby. 2024. AI-Assisted SQL Authoring at Industry Scale. arXiv:2407.13280 [cs.SE] https://arxiv.org/abs/2...
2024 arXiv
-
[162]
Srijoni Majumdar, Edith Elkind, and Evangelos Pournaras. 2025. Generative AI Voting: Fair Collective Choice is Resilient to LLM Biases and Inconsistencies. arXiv:2406.11871 [cs.AI] https://arxiv.org/abs/2406.11871
2025
-
[163]
Doan,Nam V.Nguyen, QuangPham, andNghiD
DungNguyenManh, ThangPhanChau, NamLeHai, ThongT. Doan,Nam V.Nguyen, QuangPham, andNghiD. Q.Bui
-
[164]
Manus. 2025. Manus. https://manus.im/
2025
-
[165]
Rohin Manvi, Samar Khanna, Gengchen Mai, Marshall Burke, David Lobell, and Stefano Ermon. 2024. GeoLLM: Extracting Geospatial Knowledge from Large Language Models. arXiv:2310.06213 [cs.CL] https://arxiv.org/abs/2310. 06213
2024 arXiv
-
[166]
Markowitz
David M. Markowitz. 2024. From Complexity to Clarity: How AI Enhances Perceptions of Scientists and the Public’s Understanding of Science. arXiv:2405.00706 [cs.CL] https://arxiv.org/abs/2405.00706
2024 arXiv
-
[167]
Jonathan Mast. 2025. ChatGPT’s Deep Research vs. Google’s Gemini 1.5 Pro with Deep Research: A Detailed Comparison. https://whitebeardstrategies.com/ai-prompt-engineering/chatgpts-deep-research-vs-googles-gemini-1-5- pro-with-deep-research-a-detailed-comparison/
2025
-
[168]
Mastra-AI. 2025. Mastra. https://github.com/mastra-ai/mastra
2025
-
[169]
Shray Mathur, Noah van der Vleuten, Kevin Yager, and Esther Tsai. 2024. VISION: A Modular AI Assistant for Natural Human-Instrument Interaction at Scientific User Facilities. arXiv:2412.18161 [cs.AI] https://arxiv.org/abs/2412.18161
2024 arXiv
-
[170]
Gianmarco Mengaldo. 2025. Explain the Black Box for the Sake of Science: the Scientific Method in the Era of Generative Artificial Intelligence. arXiv:2406.10557 [cs.AI] https://arxiv.org/abs/2406.10557
2025 arXiv
-
[171]
2025.MGX.dev
MGX Technologies. 2025.MGX.dev. https://mgx.dev
2025
-
[172]
Gregoire Mialon, Clementine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023. GAIA:A Benchmark for General AI Assistants. https://huggingface.co/gaia-benchmark. https://arxiv.org/pdf/2311.12983
2023 arXiv
-
[173]
Microsoft. 2023. Microsoft Copilot. https://www.microsoft.com/en-us/microsoft-copilot/organizations
2023
-
[174]
Microsoft. 2023. Semantic-kernel. https://github.com/microsoft/semantic-kernel
2023
-
[175]
mirayayerdem. 2022. Github-Copilot-Amazon-Whisperer-ChatGPT. https://github.com/mirayayerdem/Github-Copilot- Amazon-Whisperer-ChatGPT
2022
-
[176]
Mlc-ai. 2023. web-llm. https://github.com/mlc-ai/web-llm
2023
-
[177]
ModelTC. 2025. lightllm. https://github.com/ModelTC/lightllm
2025
-
[178]
Devam Mondal and Atharva Inamdar. 2024. SeqMate: A Novel Large Language Model Pipeline for Automating RNA Sequencing. arXiv:2407.03381 [q-bio.GN] https://arxiv.org/abs/2407.03381
2024 arXiv
-
[179]
Peya Mowar, Yi-Hao Peng, Jason Wu, Aaron Steinfeld, and Jeffrey P. Bigham. 2025. CodeA11y: Making AI Coding Assistants Useful for Accessible Web Development. arXiv:2502.10884 [cs.HC] https://arxiv.org/abs/2502.10884
2025 arXiv
-
[180]
Dilxat Muhtar, Zhenshi Li, Feng Gu, Xueliang Zhang, and Pengfeng Xiao. 2024. LHRS-Bot: Empowering Re- mote Sensing with VGI-Enhanced Large Multimodal Language Model. https://github.com/NJU-LHRS/LHRS-Bot. arXiv:2402.02544 [cs.CV] https://arxiv.org/abs/2402.02544
2024 arXiv
-
[181]
Manisha Mukherjee, Sungchul Kim, Xiang Chen, Dan Luo, Tong Yu, and Tung Mai. 2025. From Documents to Dialogue: Building KG-RAG Enhanced AI Assistants. arXiv:2502.15237 [cs.IR] https://arxiv.org/abs/2502.15237
2025 arXiv
-
[182]
Sheshera Mysore, Mahmood Jasim, Haoru Song, Sarah Akbar, Andre Kenneth Chase Randall, and Narges Mahyar
-
[183]
n8n. 2023. n8n. https://github.com/n8n-io/n8n
2023
-
[184]
Nanobrowser Team. 2024. Nanobrowser. https://github.com/nanobrowser/nanobrowser
2024
-
[185]
Nathalia Nascimento, Everton Guimaraes, Sai Sanjna Chintakunta, and Santhosh Anitha Boominathan. 2024. LLM4DS: Evaluating Large Language Models for Data Science Code Generation. https://github.com/DataForScience/LLM4DS. arXiv:2411.11908 [cs.SE] https://arxiv.org/abs/2411.11908
2024 arXiv
-
[186]
InProceedings of the 2023 Conference on Human Information Interaction and Retrieval (CHIIR ’23)
How Data Scientists Review the Scholarly Literature. InProceedings of the 2023 Conference on Human Information Interaction and Retrieval (CHIIR ’23). ACM, 137–152. doi:10.1145/3576840.3578309
2023
-
[187]
Alex Nguyen, Zilong Wang, Jingbo Shang, and Dheeraj Mekala. 2024. DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering. arXiv:2404.00439 [cs.CL] https://arxiv.org/abs/2404.00439
2024 arXiv
-
[188]
Nguyen, Fengchun Qiao, Arthur Trembanis, and Xi Peng
Kien X. Nguyen, Fengchun Qiao, Arthur Trembanis, and Xi Peng. 2024. SeafloorAI: A Large-scale Vision-Language Dataset for Seafloor Geological Survey. https://github.com/deep-real/SeafloorAI. arXiv:2411.00172 [cs.CV] https: //arxiv.org/abs/2411.00172
2024 arXiv
-
[189]
Ziqi Ni, Yahao Li, Kaijia Hu, Kunyuan Han, Ming Xu, Xingyu Chen, Fengqi Liu, Yicong Ye, and Shuxin Bai
-
[190]
Khanh Nghiem, Anh Minh Nguyen, and Nghi D. Q. Bui. 2024. Envisioning the Next-Generation AI Coding Assistants: Insights & Proposals. arXiv:2403.14592 [cs.SE] https://arxiv.org/abs/2403.14592 A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications 89
2024 arXiv
-
[191]
Kanda, and Haruka Ozaki
Koji Ochiai, Yuya Tahara-Arai, Akari Kato, Kazunari Kaizu, Hirokazu Kariyazaki, Makoto Umeno, Koichi Takahashi, Genki N. Kanda, and Haruka Ozaki. 2025. Automating Care by Self-maintainability for Full Laboratory Automation. arXiv:2501.05789 [q-bio.QM] https://arxiv.org/abs/2501.05789
2025 arXiv
-
[192]
Ollama. 2023. Ollama. https://github.com/ollama/ollama
2023
-
[193]
Open Manus Team. 2025. OpenManus. https://github.com/mannaandpoem/OpenManus
2025
-
[194]
arXiv:2411.08063 [physics.soc-ph] https://arxiv.org/abs/2411.08063
MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration. arXiv:2411.08063 [physics.soc-ph] https://arxiv.org/abs/2411.08063
-
[195]
Alexander Novikov, Ngân V˜ u, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaud- huri, George Holland, Alex Davies, Sebastian Nowozin, Pu...
2025
-
[196]
OpenAI. 2025. Deep Research System Card. https://cdn.openai.com/deep-research-system-card.pdf
2025
-
[198]
OpenAI. 2025. Introducing OpenAI o3 and o4-mini. https://openai.com/index/introducing-o3-and-o4-mini/
2025
-
[199]
OpenAI. 2025. codex. https://github.com/openai/codex
2025
-
[200]
OpenAI. 2025. Compare models - OpenAI API. https://platform.openai.com/docs/models/compare?model=o3
2025
-
[201]
OpenAI. 2025. Thinking with images. https://openai.com/index/thinking-with-images/
2025
-
[202]
OpenBMB. 2023. XAgent. https://github.com/OpenBMB/XAgent
2023
-
[203]
Orkes. 2022. Orkes. https://orkes.io/use-cases/agentic-workflows
2022
-
[204]
OpenAI. 2025. OpenAI Agents SDK. https://github.com/openai/openai-agents-python
2025
-
[205]
OpenAI. 2025. OpenAI o3 and o4-mini System Card. https://cdn.openai.com/pdf/2221c875-02dc-4789-800b- e7758f3722c1/o3-and-o4-mini-system-card.pdf
2025
-
[206]
Mahabubur Rahman, and Mst
Md Sultanul Islam Ovi, Nafisa Anjum, Tasmina Haque Bithe, Md. Mahabubur Rahman, and Mst. Shahnaj Akter Smrity
-
[207]
Carlos Alves Pereira, Tanay Komarlu, and Wael Mobeirek. 2023. The Future of AI-Assisted Writing. arXiv:2306.16641 [cs.HC] https://arxiv.org/abs/2306.16641
2023 arXiv
-
[208]
Mike Perkins and Jasper Roe. 2024. Generative AI Tools in Academic Research: Applications and Implications for Qualitative and Quantitative Research Methodologies. arXiv:2408.06872 [cs.HC] https://arxiv.org/abs/2408.06872
2024 arXiv
-
[209]
Takauki Osogami. 2025. Position: AI agents should be regulated based on autonomous action sequences. arXiv:2503.04750 [cs.CY] https://arxiv.org/abs/2503.04750
2025 arXiv
-
[210]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022 arXiv
-
[211]
Tomas Petricek, Gerrit J. J. van den Burg, Alfredo Nazábal, Taha Ceritli, Ernesto Jiménez-Ruiz, and Christopher K. I. Williams. 2022. AI Assistants: A Framework for Semi-Automated Data Wrangling. arXiv:2211.00192 [cs.DB] https://arxiv.org/abs/2211.00192
2022 arXiv
-
[212]
arXiv:2409.19922 [cs.SE] https://arxiv.org/abs/2409.19922
Benchmarking ChatGPT, Codeium, and GitHub Copilot: A Comparative Study of AI-Driven Programming and Debugging Assistants. arXiv:2409.19922 [cs.SE] https://arxiv.org/abs/2409.19922
-
[213]
Evangelos Pournaras. 2023. Science in the Era of ChatGPT, Large Language Models and Generative AI: Challenges for Research Ethics and How to Respond. arXiv:2305.15299 [cs.CY] https://arxiv.org/abs/2305.15299
2023 arXiv
-
[214]
Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024. Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented 90 Xu et al. Generation Track. arXiv:2406.16828 [cs.IR] https://...
2024 arXiv
-
[215]
Perplexity. 2025. Introducing Perplexity Deep Research. https://www.perplexity.ai/hub/blog/introducing-perplexity- deep-research
2025
-
[216]
Perplexity. 2025. Sonar by Perplexity. https://docs.perplexity.ai/guides/model-cards#research-models
2025
-
[217]
Pythagora-io. 2024. gpt-pilot. https://github.com/Pythagora-io/gpt-pilot
2024
-
[218]
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. 2025. Humanity’s Last Exam. arXiv:2501.14249 [cs.LG] https://arxiv.org/abs/ 2501.14249
2025 arXiv
-
[219]
Laryn Qi, J. D. Zamfirescu-Pereira, Taehan Kim, Björn Hartmann, John DeNero, and Narges Norouzi. 2024. A Knowledge-Component-Based Methodology for Evaluating AI Assistants. arXiv:2406.05603 [cs.CY] https://arxiv.org/ abs/2406.05603
2024 arXiv
-
[220]
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024. ChatDev: Communicative Agents for Software Development. https://github.com/OpenBMB/ChatDev. https://aclant...
2024
-
[221]
Reeves, Jaromir Savelka, David H
James Prather, Juho Leinonen, Natalie Kiesler, Jamie Gorson Benario, Sam Lau, Stephen MacNeil, Narges Norouzi, Simone Opel, Vee Pettit, Leo Porter, Brent N. Reeves, Jaromir Savelka, David H. Smith IV, Sven Strickroth, and Daniel Zingaro. 2024. Beyond the Hype: A Comprehensive ...
2024 arXiv
-
[222]
Pydantic. 2024. Pydantic-AI. https://github.com/pydantic/pydantic-ai
2024
-
[223]
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023. ToolLLM: Facilitating Large Language Model...
2023 arXiv
-
[224]
Jingyuan Qi, Zian Jia, Minqian Liu, Wangzhi Zhan, Junkai Zhang, Xiaofei Wen, Jingru Gan, Jianpeng Chen, Qin Liu, Mingyu Derek Ma, Bangzheng Li, Haohui Wang, Adithya Kulkarni, Muhao Chen, Dawei Zhou, Ling Li, Wei Wang, and Lifu Huang. 2024. MetaScientist: A Human-AI Synergistic...
2024 arXiv
-
[225]
Joaquin Ramirez-Medina, Mohammadmehdi Ataei, and Alidad Amirfazli. 2025. Accelerating Scientific Research Through a Multi-LLM Framework. arXiv:2502.07960 [physics.app-ph] https://arxiv.org/abs/2502.07960
2025 arXiv
-
[226]
Ruchit Rawal, Victor-Alexandru Pădurean, Sven Apel, Adish Singla, and Mariya Toneva. 2024. Hints Help Finding and Fixing Bugs Differently in Python and Text-based Program Representations. arXiv:2412.12471 [cs.SE] https: //arxiv.org/abs/2412.12471
2024 arXiv
-
[227]
Shuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang, Ningyu Zhang, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen. 2025. Benchmarking Agentic Workflow Generation. https://github.com/zjunlp/WorfBench. arXiv:2410.07869 [cs.CL] https://arxiv.org/abs/2410.07869
2025 arXiv
-
[229]
Restate. 2024. Restate. https://restate.dev/
2024
-
[230]
Qwen LM. 2024. Qwen-Agent. https://github.com/QwenLM/Qwen-Agent
2024
-
[231]
Filippo Ricca, Alessandro Marchetto, and Andrea Stocco. 2025. A Multi-Year Grey Literature Review on AI-assisted Test Automation. https://arxiv.org/pdf/2408.06224
2025 arXiv
-
[232]
Nathalie Riche, Anna Offenwanger, Frederic Gmeiner, David Brown, Hugo Romat, Michel Pahud, Nicolai Marquardt, Kori Inkpen, and Ken Hinckley. 2025. AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools. http...
2025 arXiv
-
[233]
Runtao Ren, Jian Ma, and Jianxi Luo. 2025. Large language model for patent concept generation.Advanced Engineering Informatics 65 (May 2025), 103301. doi:10.1016/j.aei.2025.103301
2025
-
[234]
ResearchRabbit. 2025. ResearchRabbit. https://www.researchrabbit.ai/
2025
-
[235]
Run-llama. 2023. LlamaIndex. https://github.com/run-llama/llama_index
2023
-
[236]
reworkd. 2023. AgentGPT. https://github.com/reworkd/AgentGPT
2023
-
[237]
SamuelSchmidgall. 2025. AgentLaboratory. https://github.com/SamuelSchmidgall/AgentLaboratory. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications 91
2025
-
[238]
Huberman
Thomas Sandholm, Sarah Dong, Sayandev Mukherjee, John Feland, and Bernardo A. Huberman. 2024. Semantic Navigation for AI-assisted Ideation. arXiv:2411.03575 [cs.HC] https://arxiv.org/abs/2411.03575
2024 arXiv
-
[239]
Lavista Ferres
Anthony Cintron Roman, Jennifer Wortman Vaughan, Valerie See, Steph Ballard, Jehu Torres, Caleb Robinson, and Juan M. Lavista Ferres. 2024. Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments. arXiv:2312.06153 [cs.LG] https://arxiv....
2024 arXiv
-
[240]
Kaushik Roy, Vedant Khandelwal, Harshul Surana, Valerie Vera, Amit Sheth, and Heather Heckman. 2023. GEAR-Up: Generative AI and External Knowledge-based Retrieval Upgrading Scholarly Article Searches for Systematic Reviews. arXiv:2312.09948 [cs.IR] https://arxiv.org/abs/2312.09948
2023 arXiv
-
[241]
Schuemie, M
Martijn J. Schuemie, M. Soledad Cepeda, Marc A. Suchard, Jianxiao Yang, Yuxi Tian, Alejandro Schuler, Patrick B. Ryan, David Madigan, and George Hripcsak. 2020. How Confident Are We About Observational Findings in Healthcare: A Benchmark Study.Harvard Data Science Review2, 1 (...
2020 doi
-
[242]
Sergey V Samsonau, Aziza Kurbonova, Lu Jiang, Hazem Lashen, Jiamu Bai, Theresa Merchant, Ruoxi Wang, Laiba Mehnaz, Zecheng Wang, and Ishita Patil. 2024. Artificial Intelligence for Scientific Research: Authentic Research Education Framework. arXiv:2210.08966 [cs.CY] https://ar...
2024 arXiv
-
[243]
Scite. 2025. Scite. https://scite.ai/
2025
-
[244]
Agnia Sergeyuk, Yaroslav Golubev, Timofey Bryksin, and Iftekhar Ahmed. 2025. Using AI-based coding assistants in practice: State of affairs, perceptions, and ways forward.Information and Software Technology178 (Feb. 2025), 107610. doi:10.1016/j.infsof.2024.107610
2025
-
[245]
Lindsay Sanneman and Julie Shah. 2021. Explaining Reward Functions to Humans for Better Human-Robot Collabo- ration. arXiv:2110.04192 [cs.RO] https://arxiv.org/abs/2110.04192
2021 arXiv
-
[246]
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum. 2025. Agent Laboratory: Using LLM Agents as Research Assistants. arXiv:2501.04227 [cs.HC] https://arxiv.org/abs/2501.04227
2025 arXiv
-
[247]
Zejiang Shen, Tal August, Pao Siangliulue, Kyle Lo, Jonathan Bragg, Jeff Hammerbacher, Doug Downey, Joseph Chee Chang, and David Sontag. 2023. Beyond Summarization: Designing AI Support for Real-World Expository Writing Tasks. arXiv:2304.02623 [cs.CL] https://arxiv.org/abs/2304.02623
2023 arXiv
-
[248]
Scispace. 2024. Scispace. https://scispace.com/
2024
-
[249]
Michael Shumer. 2025. OpenDeepResearcher. https://github.com/mshumer/OpenDeepResearcher
2025
-
[250]
Significant-Gravitas. 2023. AutoGPT. https://github.com/Significant-Gravitas/AutoGPT
2023
-
[251]
Mahsa Shamsabadi and Jennifer D’Souza. 2024. A FAIR and Free Prompt-based Research Assistant. arXiv:2405.14601 [cs.CL] https://arxiv.org/abs/2405.14601
2024 arXiv
-
[252]
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. https://github.com/microsoft/JARVIS. https: //arxiv.org/pdf/2303.17580
2023 arXiv
-
[253]
Michael Skarlinski, Tyler Nadolski, James Braza, Remo Storni, Mayk Caldas, Ludovico Mitchener, Michaela Hinks, Andrew White, and Sam Rodriques. 2025. FutureHouse Platform: Superintelligent AI Agents for Scientific Discovery. https://www.futurehouse.org/research-announcements/l...
2025
-
[254]
Shuming Shi, Enbo Zhao, Duyu Tang, Yan Wang, Piji Li, Wei Bi, Haiyun Jiang, Guoping Huang, Leyang Cui, Xinting Huang, Cong Zhou, Yong Dai, and Dongyang Ma. 2022. Effidit: Your AI Writing Assistant. arXiv:2208.01815 [cs.CL] https://arxiv.org/abs/2208.01815
2022 arXiv
-
[255]
Jamshid Sourati and James Evans. 2021. Accelerating science with human versus alien artificial intelligences. arXiv:2104.05188 [cs.AI] https://arxiv.org/abs/2104.05188
2021 arXiv
-
[256]
Jamshid Sourati and James Evans. 2023. Accelerating science with human-aware artificial intelligence. arXiv:2306.01495 [cs.AI] https://arxiv.org/abs/2306.01495
2023 arXiv
-
[257]
David Silver and Richard Sutton. 2025. Welcome to the Era of Experience. https://storage.googleapis.com/deepmind- media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf
2025
-
[258]
It is there, and you need it, so why do you not use it?
Auste Simkute, Ewa Luger, Michael Evans, and Rhianne Jones. 2024. "It is there, and you need it, so why do you not use it?" Achieving better adoption of AI systems by domain experts, in the case study of natural science research. arXiv:2403.16895 [cs.HC] https://arxiv.org/abs/...
2024 arXiv
-
[259]
Suzhou Yuling Artificial Intelligence Technology Co
Ltd. Suzhou Yuling Artificial Intelligence Technology Co. 2023. Dify: Open-source LLM Application Development Platform. https://dify.ai/
2023
-
[260]
Clark, Hao He, Haoran He, Jie Min, Xinlei Zhang, Simin Zheng, Zhiyang Zhang, Xinwei Deng, and Yili Hong
Xinyi Song, Kexin Xie, Lina Lee, Ruizhe Chen, Jared M. Clark, Hao He, Haoran He, Jie Min, Xinlei Zhang, Simin Zheng, Zhiyang Zhang, Xinwei Deng, and Yili Hong. 2025. Performance Evaluation of Large Language Models in Statistical Programming. arXiv:2502.13117 [stat.AP] https://...
2025 arXiv
-
[261]
Brian Tang and Kang G. Shin. 2024. Steward: Natural Language Web Automation. arXiv:2409.15441 [cs.AI] https://arxiv.org/abs/2409.15441
2024 arXiv
-
[262]
Jiabin Tang, Tianyu Fan, and Chao Huang. 2025. AutoAgent: A Fully-Automated and Zero-Code Framework for LLM Agents. arXiv:2502.05957 [cs.AI] https://arxiv.org/abs/2502.05957
2025
-
[263]
StanfordNLP. 2024. DSPy. https://github.com/stanfordnlp/dspy
2024
-
[264]
Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, and Nanqing Dong. 2025. Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System. arXiv:24...
2025 arXiv
-
[265]
Temporalio. 2020. Temporal. https://github.com/temporalio/temporal
2020
-
[266]
Xin Tan, Xiao Long, Xianjun Ni, Yinghao Zhu, Jing Jiang, and Li Zhang. 2024. How far are AI-powered programming assistants from meeting developers’ needs? arXiv:2404.12000 [cs.SE] https://arxiv.org/abs/2404.12000
2024 arXiv
-
[267]
TheBlewish. 2024. Automated-AI-Web-Researcher-Ollama. https://github.com/TheBlewish/Automated-AI-Web- Researcher-Ollama
2024
-
[268]
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, Shengyan Liu, Di Luo, Yutao Ma, Hao Tong, Kha Trinh, Chenyu Tian, Zihan Wang, Bohao Wu, Yanyu Xiong, Shengzhu Yin, Minhui Zhu, Kilian Lieret, Yanx...
2024 arXiv
-
[269]
Yan Tang. 2025. deep_research_agent. https://github.com/grapeot/deep_research_agent. 92 Xu et al
2025
-
[270]
Tadahiro Taniguchi, Shiro Takagi, Jun Otsuka, Yusuke Hayashi, and Hiro Taiyo Hamada. 2024. Collective Predictive Coding as Model of Science: Formalizing Scientific Activities Towards Generative Science. arXiv:2409.00102 [physics.soc- ph] https://arxiv.org/abs/2409.00102
2024 arXiv
-
[271]
Benjamin Towle and Ke Zhou. 2024. Enhancing AI Assisted Writing with One-Shot Implicit Negative Feedback. arXiv:2410.11009 [cs.CL] https://arxiv.org/abs/2410.11009
2024 arXiv
-
[272]
Enkeleda Thaqi, Mohamed Omar Mantawy, and Enkelejda Kasneci. 2024. SARA: Smart AI Reading Assistant for Reading Comprehension. InProceedings of the 2024 Symposium on Eye Tracking Research and Applications (ETRA ’24). ACM, 1–3. doi:10.1145/3649902.3655661
2024
-
[273]
Wang, Sabrina A Sgandurra, Reza Hadi Mogavi, and Lennart E
Joseph Tu, Hilda Hadan, Derrick M. Wang, Sabrina A Sgandurra, Reza Hadi Mogavi, and Lennart E. Nacke. 2024. Augmenting the Author: Exploring the Potential of AI Collaboration in Academic Writing. arXiv:2404.16071 [cs.HC] https://arxiv.org/abs/2404.16071
2024 arXiv
-
[274]
Su, and Linjun Zhang
Xinming Tu, James Zou, Weijie J. Su, and Linjun Zhang. 2023. What Should Data Science Education Do with Large Language Models? arXiv:2307.02792 [cs.CY] https://arxiv.org/abs/2307.02792
2023 arXiv
-
[275]
Tiukova, Daniel Brunnsåker, Erik Y
Ievgeniia A. Tiukova, Daniel Brunnsåker, Erik Y. Bjurström, Alexander H. Gower, Filip Kronström, Gabriel K. Reder, Ronald S. Reiserer, Konstantin Korovin, Larisa B. Soldatova, John P. Wikswo, and Ross D. King. 2024. Genesis: Towards the Automation of Systems Biology Research. ...
2024 arXiv
-
[276]
Irina Tolstykh, Aleksandra Tsybina, Sergey Yakubson, Aleksandr Gordeev, Vladimir Dokholyan, and Maksim Kuprashe- vich. 2024. GigaCheck: Detecting LLM-generated Content. arXiv:2410.23728 [cs.CL] https://arxiv.org/abs/2410.23728
2024 arXiv
-
[277]
Rasmus Ulfsnes, Nils Brede Moe, Viktoria Stray, and Marianne Skarpen. 2024. Transforming Software Development with Generative AI: Empirical Insights on Collaboration and Workflow. arXiv:2405.01543 [cs.SE] https://arxiv.org/ abs/2405.01543
2024 arXiv
-
[278]
Thanh-Dat Truong, Hoang-Quan Nguyen, Xuan-Bac Nguyen, Ashley Dowling, Xin Li, and Khoa Luu. 2025. Insect- Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding. https: //uark-cviu.github.io/projects/insect-foundation/. arXiv:2502....
2025 arXiv
-
[279]
Jones, Oisin Mac Aodha, Sara Beery, and Grant Van Horn
Edward Vendrow, Omiros Pantazis, Alexander Shepard, Gabriel Brostow, Kate E. Jones, Oisin Mac Aodha, Sara Beery, and Grant Van Horn. 2024. INQUIRE: A Natural World Text-to-Image Retrieval Benchmark. https://inquire- benchmark.github.io/. arXiv:2411.02537 [cs.CV] https://arxiv....
2024 arXiv
-
[280]
Vercel. 2020. Vercel. https://vercel.com/
2020
-
[281]
Michele Tufano, Anisha Agarwal, Jinu Jang, Roshanak Zilouchian Moghaddam, and Neel Sundaresan. 2024. AutoDev: Automated AI-Driven Development. arXiv:2403.08299 [cs.SE] https://arxiv.org/abs/2403.08299
2024 arXiv
-
[282]
Aleksei Turobov, Diane Coyle, and Verity Harding. 2024. Using ChatGPT for Thematic Analysis. arXiv:2405.08828 [cs.HC] https://arxiv.org/abs/2405.08828
2024 arXiv
-
[283]
Weisz, Xuye Liu, Lingfei Wu, and Casey Dugan
April Yi Wang, Dakuo Wang, Jaimie Drozdal, Michael Muller, Soya Park, Justin D. Weisz, Xuye Liu, Lingfei Wu, and Casey Dugan. 2022. Documentation Matters: Human-Centered AI System to Assist Data Science Code Documentation in Computational Notebooks.ACM Transactions on Computer...
2022 doi
-
[284]
Stanford University. 2025. STORM. https://storm.genie.stanford.edu/
2025
-
[285]
Suyuan Wang, Xueqian Yin, Menghao Wang, Ruofeng Guo, and Kai Nan. 2024. EvoPat: A Multi-LLM-based Patents Summarization and Analysis Agent. arXiv:2412.18100 [cs.DL] https://arxiv.org/abs/2412.18100
2024 arXiv
-
[286]
Tiannan Wang, Jiamin Chen, Qingrui Jia, Shuai Wang, Ruoyu Fang, Huilin Wang, Zhaowei Gao, Chunzhao Xie, Chuou Xu, Jihong Dai, Yibin Liu, Jialong Wu, Shengwei Ding, Long Li, Zhiwei Huang, Xinle Deng, Teng Yu, Gangan Ma, Han A Comprehensive Survey of Deep Research: Systems, Meth...
2024 arXiv
-
[287]
Vllm-project. 2023. vllm. https://github.com/vllm-project/vllm
2023
-
[288]
Thiemo Wambsganss, Xiaotian Su, Vinitra Swamy, Seyed Parsa Neshaei, Roman Rietsche, and Tanja Käser. 2023. Unraveling Downstream Gender Bias from Large Language Models: A Study on AI Educational Writing Assistance. arXiv:2311.03311 [cs.CL] https://arxiv.org/abs/2311.03311
2023 arXiv
-
[289]
Ying-Mei Wang and Tzeng-J Chen. 2025. AI’s deep research revolution: Transforming biomedical literature analysis. https://journals.lww.com/jcma/citation/9900/ai_s_deep_research_revolution__transforming.508.aspx
2025
-
[290]
Vera Liao, Yunfeng Zhang, Udayan Khurana, Horst Samulowitz, Soya Park, Michael Muller, and Lisa Amini
Dakuo Wang, Q. Vera Liao, Yunfeng Zhang, Udayan Khurana, Horst Samulowitz, Soya Park, Michael Muller, and Lisa Amini. 2021. How Much Automation Does a Data Scientist Want? arXiv:2101.03970 [cs.LG] https://arxiv.org/abs/ 2101.03970
2021 arXiv
-
[291]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903 [cs.CL] https://arxiv.org/abs/2201.11903
2023 arXiv
-
[292]
Shufa Wei, Xiaolong Xu, Xianbiao Qi, Xi Yin, Jun Xia, Jingyi Ren, Peijun Tang, Yuxiang Zhong, Yihao Chen, Xiaoqin Ren, Yuxin Liang, Liankai Huang, Kai Xie, Weikang Gui, Wei Tan, Shuanglong Sun, Yongquan Hu, Qinxian Liu, Nanjin Li, Chihao Dai, Lihua Wang, Xiaohui Liu, Lei Zhang...
2023 arXiv
-
[293]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171 [cs.CL] https://arxiv.org/abs/2203.11171
2023 arXiv
-
[294]
Yao Wang, Mingxuan Cui, and Arthur Jiang. 2025. Enabling AI Scientists to Recognize Innovation: A Domain-Agnostic Algorithm for Assessing Novelty. arXiv:2503.01508 [cs.AI] https://arxiv.org/abs/2503.01508
2025
-
[296]
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024. Measuring short-form factuality in large language models. https://cdn.openai.com/papers/ simpleqa.pdf
2024
-
[300]
Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. 2025. CycleResearcher: Improving Automated Research via Automated Review. arXiv:2411.00816 [cs.CL] https://arxiv.org/ abs/2411.00816
2025 arXiv
-
[2023]
arXiv:2307.07049 [cs.CL] https://arxiv.org/abs/2307.07049 82 Xu et al
MegaWika: Millions of reports and their sources across 50 diverse languages. arXiv:2307.07049 [cs.CL] https://arxiv.org/abs/2307.07049 82 Xu et al
-
[2024]
https: //uni-medical.github.io/GMAI-MMBench.github.io/
GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI. https: //uni-medical.github.io/GMAI-MMBench.github.io/. arXiv:2408.03361 [eess.IV] https://arxiv.org/abs/2408.03361
-
[2025]
https://github.com/FSoft-AI4Code/CodeMMLU
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs. https://github.com/FSoft-AI4Code/CodeMMLU. arXiv:2410.01999 [cs.SE] https://arxiv.org/abs/2410.01999
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.