Pith. sign in

REVIEW 11 cited by

Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05688 v1 pith:D5JZ5ARA submitted 2024-06-09 cs.CL cs.AIcs.LG

Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions

classification cs.CL cs.AIcs.LG
keywords peer-reviewprocessapplicationsdatasetllmsmulti-turnpeerreview
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models (LLMs) have demonstrated wide-ranging applications across various fields and have shown significant potential in the academic peer-review process. However, existing applications are primarily limited to static review generation based on submitted papers, which fail to capture the dynamic and iterative nature of real-world peer reviews. In this paper, we reformulate the peer-review process as a multi-turn, long-context dialogue, incorporating distinct roles for authors, reviewers, and decision makers. We construct a comprehensive dataset containing over 26,841 papers with 92,017 reviews collected from multiple sources, including the top-tier conference and prestigious journal. This dataset is meticulously designed to facilitate the applications of LLMs for multi-turn dialogues, effectively simulating the complete peer-review process. Furthermore, we propose a series of metrics to evaluate the performance of LLMs for each role under this reformulated peer-review setting, ensuring fair and comprehensive evaluations. We believe this work provides a promising perspective on enhancing the LLM-driven peer-review process by incorporating dynamic, role-based interactions. It aligns closely with the iterative and interactive nature of real-world academic peer review, offering a robust foundation for future research and development in this area. We open-source the dataset at https://github.com/chengtan9907/ReviewMT.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents

    cs.CL 2026-04 unverdicted novelty 7.0

    ReviewGrounder decomposes review generation into rubric-guided drafting and tool-integrated grounding stages, outperforming larger baseline models on a new benchmark measuring alignment with human judgments and review...

  2. Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review

    cs.CL 2026-01 unverdicted novelty 7.0

    Introduces Re3Align dataset of review-response-revision triplets, REspGen author-in-the-loop generation framework, and REspEval multi-metric suite for controllable peer-review response generation.

  3. Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review

    cs.CL 2026-01 conditional novelty 7.0

    An author-in-the-loop framework with a new 15k review-response-revision dataset and 20+ metrics shows that supplying author revision signals improves LLM-generated peer-review rebuttals.

  4. PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

    cs.AI 2026-05 unverdicted novelty 6.0

    PRAIB reveals LLM reviews are less variable, positively biased, overconfident, longer, and overlook atomic weaknesses noted by humans compared to real reviewer feedback.

  5. GraphReview: Scientific Paper Evaluation via LLM-Based Graph Message Passing

    cs.CL 2026-05 unverdicted novelty 6.0

    GraphReview models paper evaluation as LLM-driven message passing on a semantic paper graph that links intrinsic quality, contemporaneous papers, and prior work, then applies Personalized PageRank for ranking and revi...

  6. EGTR-Review: Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation

    cs.CL 2026-06 unverdicted novelty 5.0

    EGTR-Review distills a multi-agent evidence-grounded review generator into an efficient student model that outperforms baselines on quality, grounding, and traceability while using fewer tokens.

  7. SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

    cs.CL 2026-04 unverdicted novelty 5.0

    SafeReview trains a Generator to create adversarial prompts and a Defender to detect them via co-evolution with an IR-GAN-inspired loss, claiming better resilience than static defenses for LLM-based peer review.

  8. Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI

    cs.CL 2026-04 unverdicted novelty 5.0

    Peer review reports in AI conferences have grown longer and more standardized after LLMs, with increased emphasis on surface-level clarity and summaries at the expense of deeper critiques on originality and replicability.

  9. LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

    cs.CL 2026-06 unverdicted novelty 4.0

    A survey synthesizing LLM methods for peer review critique generation and score prediction, including taxonomies, benchmark limitations, domain biases, and robustness risks such as prompt injection.

  10. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 conditional novelty 4.0

    AI can generate research artifacts faster than it can verify them, so across all eight lifecycle stages the credible deployment mode is human-governed collaboration rather than full autonomy.

  11. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.