Pith. sign in

REVIEW 11 cited by

DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08569 v1 pith:RNLCU7BB submitted 2025-03-11 cs.CL cs.LG

classification cs.CLcs.LG
keywords reviewllm-basedstructureddatasetdeepreviewdeepreviewer-14bachievesaddress
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) are increasingly utilized in scientific research assessment, particularly in automated paper review. However, existing LLM-based review systems face significant challenges, including limited domain expertise, hallucinated reasoning, and a lack of structured evaluation. To address these limitations, we introduce DeepReview, a multi-stage framework designed to emulate expert reviewers by incorporating structured analysis, literature retrieval, and evidence-based argumentation. Using DeepReview-13K, a curated dataset with structured annotations, we train DeepReviewer-14B, which outperforms CycleReviewer-70B with fewer tokens. In its best mode, DeepReviewer-14B achieves win rates of 88.21\% and 80.20\% against GPT-o1 and DeepSeek-R1 in evaluations. Our work sets a new benchmark for LLM-based paper review, with all resources publicly available. The code, model, dataset and demo have be released in http://ai-researcher.net.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas

    cs.CL 2025-06 conditional novelty 8.0 of 10

    A randomized execution study with 43 experts shows that LLM-generated research ideas lose more of their appeal than human ideas when actually implemented, reversing part of their ideation-stage advantage.

  2. OmniPresent: Generating Coherent Presentation Suites from Scientific Papers

    cs.SE 2026-07 conditional novelty 6.5 of 10

    A multi-agent HTML pipeline with shared knowledge and cross-artifact verify-and-repair generates coherent poster/slides/video/page suites from papers and beats specialized baselines on OmniPreBench.

  3. Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance

    cs.AI 2026-01 conditional novelty 6.0 of 10

    A multi-agent 'verify-then-write' system for writing peer-review rebuttals beats direct LLM prompting on a new benchmark, but the gains are measured by an LLM judge, not by the original reviewers.

  4. OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

    cs.CL 2025-10 reject novelty 6.0 of 10

    A tool-augmented reward model trained with GRPO on 27K synthetic pairs beats existing reward models on long-form QA judgment and improves downstream alignment.

  5. ReviewRL: Towards Automated Scientific Review with RL

    cs.CL 2025-08 conditional novelty 6.0 of 10

    ReviewRL combines arXiv retrieval, supervised fine-tuning, and reinforcement learning with a composite reward to generate paper reviews that better match human ratings and judged quality.

  6. THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning?

    cs.AI 2025-06 reject novelty 6.0 of 10

    THE-Tree constructs causally-linked semantic evolution trees from surveys and literature, and the authors report improved graph completion, future prediction, and LLM-based paper evaluation.

  7. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  8. When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review

    cs.CY 2025-09 conditional novelty 4.0 of 10

    GPT-5-mini gives weaker papers systematically higher scores than human reviewers, and hidden field-specific prompts in PDFs can force it to assign perfect scores or suppress weaknesses.

  9. How Far Are AI Scientists from Changing the World?

    cs.AI 2025-07 conditional novelty 4.0 of 10

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

  10. Deep Research Agents: A Systematic Examination And Roadmap

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey that organizes LLM-powered deep research agents into static versus dynamic workflows and single versus multi agent architectures, and reviews their benchmarks and open challenges.

  11. AI Scientists Fail Without Strong Implementation Capability

    cs.AI 2025-06 conditional novelty 4.0 of 10

    AI scientist systems can propose ideas but cannot reliably implement and verify experiments, making the implementation gap, not idea generation, the current bottleneck.

Pith tools