Pith. sign in

REVIEW 5 cited by

AutoSurvey: Large Language Models Can Automatically Write Surveys

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10252 v2 pith:EROTVZ6O submitted 2024-06-10 cs.IR cs.AIcs.CL

classification cs.IRcs.AIcs.CL
keywords autosurveychallengesevaluationsurveyautomatingcomprehensivecreationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces AutoSurvey, a speedy and well-organized methodology for automating the creation of comprehensive literature surveys in rapidly evolving fields like artificial intelligence. Traditional survey paper creation faces challenges due to the vast volume and complexity of information, prompting the need for efficient survey methods. While large language models (LLMs) offer promise in automating this process, challenges such as context window limitations, parametric knowledge constraints, and the lack of evaluation benchmarks remain. AutoSurvey addresses these challenges through a systematic approach that involves initial retrieval and outline generation, subsection drafting by specialized LLMs, integration and refinement, and rigorous evaluation and iteration. Our contributions include a comprehensive solution to the survey problem, a reliable evaluation method, and experimental validation demonstrating AutoSurvey's effectiveness.We open our resources at \url{https://github.com/AutoSurveys/AutoSurvey}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio

    cs.CL 2026-06 unverdicted novelty 8.0 of 10

    MetaSyn is a stage-level benchmark of 442 meta-analyses showing LLM agents retrieve up to 90.9% of eligible studies but include at most 52.7% in their final reports.

  2. RWGBench: Evaluating Scholarly Positioning in Related Work Generation

    cs.DL 2026-05 unverdicted novelty 7.0 of 10

    RWGBench measures related-work generation by citation choices, and shows citation-focused metrics expose failures that text-similarity and LLM-judge scores miss.

  3. Select, Read, and Write: A Multi-Agent Framework of Full-Text-based Related Work Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multi-agent reader-selector-writer framework improves generated related-work sections by reading full texts in a graph-guided order and compressing key information into shared memory.

  4. Automated Capability Discovery via Foundation Model Self-Exploration

    cs.LG 2025-02 conditional novelty 6.0 of 10

    ACD automatically generates thousands of open-ended tasks and clusters them into dozens of capability and failure categories, with LLM-vs-human scoring agreement (F1 = 0.86).

  5. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

Pith tools