Pith. sign in

REVIEW 8 cited by

Can large language models provide useful feedback on research papers? A large-scale empirical analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01783 v1 pith:NV3TWYHG submitted 2023-10-03 cs.LG cs.AIcs.CLcs.HC

classification cs.LGcs.AIcs.CLcs.HC
keywords feedbackgpt-4humanoverlapresearchersreviewersgeneratediclr
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Expert feedback lays the foundation of rigorous research. However, the rapid growth of scholarly production and intricate knowledge specialization challenge the conventional scientific feedback mechanisms. High-quality peer reviews are increasingly difficult to obtain. Researchers who are more junior or from under-resourced settings have especially hard times getting timely feedback. With the breakthrough of large language models (LLM) such as GPT-4, there is growing interest in using LLMs to generate scientific feedback on research manuscripts. However, the utility of LLM-generated feedback has not been systematically studied. To address this gap, we created an automated pipeline using GPT-4 to provide comments on the full PDFs of scientific papers. We evaluated the quality of GPT-4's feedback through two large-scale studies. We first quantitatively compared GPT-4's generated feedback with human peer reviewer feedback in 15 Nature family journals (3,096 papers in total) and the ICLR machine learning conference (1,709 papers). The overlap in the points raised by GPT-4 and by human reviewers (average overlap 30.85% for Nature journals, 39.23% for ICLR) is comparable to the overlap between two human reviewers (average overlap 28.58% for Nature journals, 35.25% for ICLR). The overlap between GPT-4 and human reviewers is larger for the weaker papers. We then conducted a prospective user study with 308 researchers from 110 US institutions in the field of AI and computational biology to understand how researchers perceive feedback generated by our GPT-4 system on their own papers. Overall, more than half (57.4%) of the users found GPT-4 generated feedback helpful/very helpful and 82.4% found it more beneficial than feedback from at least some human reviewers. While our findings show that LLM-generated feedback can help researchers, we also identify several limitations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 43 citations worldwide. Full citation record

  1. FARS: A Fully Automated Research System Deployed at Scale

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    FARS deployed at scale produced 166 AI/ML papers across 67 topics that received 282 structured human reviews indicating some review-worthy outputs alongside recurring failure modes.

  2. NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    NatureBench evaluates ten frontier AI coding agents on 90 tasks from Nature papers under web-search-disabled conditions and finds the strongest agent surpasses published SOTA on only 17.8% of tasks, succeeding mainly ...

  3. AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

    astro-ph.IM 2026-07 conditional novelty 6.0 of 10

    In a controlled test, three mid-2025 LLMs shared under 6% of literature references with physics experts, and 64% of their real references had at least one metadata error.

  4. Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review

    cs.DL 2026-04 conditional novelty 6.0 of 10

    An audit of 50,289 ICLR papers shows acceptance odds vary up to 8x across topics at equal reviewer scores, indicating scores are not comparable across research areas.

  5. Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development

    cs.CL 2025-05 reject novelty 6.0 of 10

    Co-Saving cuts token usage by roughly half in multi-agent software development by injecting learned shortcut instructions that bypass intermediate reasoning steps, while slightly improving a composite code-quality score.

  6. AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

    cs.CL 2026-07 conditional novelty 5.0 of 10

    AI-generated one-page research proposals are scored about the same as human-written ones by human reviewers, but AI reviewers favor AI-written proposals by roughly one point and detect authorship perfectly.

  7. Automated Novelty Evaluation of Academic Paper: A Collaborative Approach Integrating Human and Large Language Model Knowledge

    cs.CL 2025-07 reject novelty 5.0 of 10

    Method novelty prediction from peer-review novelty sentences and ChatGPT method summaries improves accuracy on ICLR 2022 data, but the benchmark leaks reviewer opinions into the input.

  8. From Image Captioning to Visual Storytelling

    cs.CL 2025-07 unverdicted novelty 4.0 of 10

    Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.

Pith tools