Pith. sign in

REVIEW 15 cited by

PaperQA: Retrieval-Augmented Generative Agent for Scientific Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07559 v2 pith:PJGDHFNR submitted 2023-12-08 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords scientificagentliteraturepaperqaacrossmodelsansweringfull-text
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) generalize well across language tasks, but suffer from hallucinations and uninterpretability, making it difficult to assess their accuracy without ground-truth. Retrieval-Augmented Generation (RAG) models have been proposed to reduce hallucinations and provide provenance for how an answer was generated. Applying such models to the scientific literature may enable large-scale, systematic processing of scientific knowledge. We present PaperQA, a RAG agent for answering questions over the scientific literature. PaperQA is an agent that performs information retrieval across full-text scientific articles, assesses the relevance of sources and passages, and uses RAG to provide answers. Viewing this agent as a question answering model, we find it exceeds performance of existing LLMs and LLM agents on current science QA benchmarks. To push the field closer to how humans perform research on scientific literature, we also introduce LitQA, a more complex benchmark that requires retrieval and synthesis of information from full-text scientific papers across the literature. Finally, we demonstrate PaperQA's matches expert human researchers on LitQA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 52 citations worldwide. Full citation record

  1. NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    NatureBench evaluates ten frontier AI coding agents on 90 tasks from Nature papers under web-search-disabled conditions and finds the strongest agent surpasses published SOTA on only 17.8% of tasks, succeeding mainly ...

  2. AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

    cs.AI 2026-06 conditional novelty 6.0 of 10

    AISE-Bench is a real-user-query benchmark with annotated API trajectories and grounded answers that exposes LLM agents' weak performance in multi-step academic information seeking.

  3. EXP-Bench: Can AI Conduct AI Research Experiments?

    cs.AI 2025-05 conditional novelty 6.0 of 10

    EXP-Bench is a new benchmark of 461 end-to-end AI research experiments, and leading AI agents complete fewer than 1 percent of them successfully.

  4. Single-agent or Multi-agent Systems? Why Not Both?

    cs.MA 2025-05 conditional novelty 6.0 of 10

    On 15 agentic benchmarks, the accuracy advantage of multi-agent LLM systems over single-agent systems mostly disappears with stronger base models, and a hybrid single/multi-agent cascade improves accuracy and cuts cost.

  5. Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark (TruthHypo) and a knowledge-grounded hallucination detector (KnowHD) show that grounding scores can partially select truthful LLM-generated biomedical hypotheses, but the result is at risk from knowled...

  6. VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

    cs.IR 2026-07 conditional novelty 5.0 of 10

    An agentic RAG framework that splits corpus-level discovery (vector retrieval) from within-paper evidence localization (tree navigation) reports top scores on QASPER, LitQA2, and a new MOSAIC benchmark.

  7. VASP Agent: An Agentic Framework for Autonomous First-principles Calculations

    cs.AI 2025-12 conditional novelty 5.0 of 10

    An LLM-driven agent with predefined VASP workflows and parameter-checking tools completes DFT simulation tasks more reliably and accurately than standalone LLMs, with a new 80-task benchmark.

  8. GRIP: In-Parameter Graph Reasoning through Fine-Tuning Large Language Models

    cs.CL 2025-11 reject novelty 5.0 of 10

    An LLM can memorize a knowledge graph into LoRA weights and answer relation/reasoning queries about it without graph context, but the evaluation partly trains on the test task.

  9. DeepResearch$^{\text{Eco}}$: A Recursive Agentic Workflow for Complex Scientific Question Answering in Ecology

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A recursive agentic pipeline for literature synthesis showing a 21-fold source increase and 14.9-fold density gain when depth and breadth are raised.

  10. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  11. How Far Are AI Scientists from Changing the World?

    cs.AI 2025-07 conditional novelty 4.0 of 10

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

  12. Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review

    cs.IR 2025-07 conditional novelty 4.0 of 10

    A DPO-tuned agentic hybrid RAG system that routes queries between a knowledge graph and a vector store beat a static baseline on a self-generated benchmark.

  13. Reading Between the Timelines: RAG for Answering Diachronic Questions

    cs.CL 2025-07 conditional novelty 4.0 of 10

    TA-RAG uses LLM-extracted time intervals, time-filtered retrieval with averaged temporal query embeddings, and chronologically ordered context to beat standard RAG by 13-27 points on the new ADQAB benchmark of 525 mul...

  14. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

  15. On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

    cs.IR 2025-06 conditional novelty 4.0 of 10

    Preprocessing financial PDFs into text, tables, and chart data with existing tools improves LLM question-answering accuracy over direct GPT-4o image input in a small private evaluation.

Pith tools