Pith. sign in

REVIEW 3 cited by

BotChat: Evaluating LLMs' Capabilities of Having Multi-Turn Dialogues

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.13650 v1 pith:MZTLOKP7 submitted 2023-10-20 cs.CL

classification cs.CL
keywords dialoguesmulti-turnllmsgeneratecapabilityevaluationgpt-4human
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Interacting with human via high-quality multi-turn dialogues is a key feature of large language models (LLMs). However, human-based evaluation of such capability involves intensive manual labor. This report provides a preliminary evaluation of existing large language models for human-style multi-turn chatting, through an LLM-based approach. We start from real-world human dialogues and keep the very first utterances as the ChatSEED. Then we prompt LLMs to generate a full multi-turn dialogue (tens of utterances) based on the ChatSEED, utterance by utterance. Finally, we adopt state-of-the-art LLMs (GPT-4, \etc) as the judge to evaluate the generated dialogues. With different evaluation protocols, we come to substantially identical conclusions. We find that GPT-4 can generate human-style multi-turn dialogues with impressive quality, significantly outperforms its counterparts. It's difficult for a discriminator to distinguish between GPT-4 generated dialogues and human dialogues. In contrast, other LLMs struggle to generate multi-turn dialogues of satisfactory quality due to poor instruction-following capability, tendency to generate lengthy utterances, or limited general capability. All data and codes will be provided in https://github.com/open-compass/BotChat/ and we hope they can serve as a valuable resource for evaluating multi-turn chatting capabilities of LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models

    cs.HC 2025-08 reject novelty 5.0 of 10

    OnGoal is an LLM chat interface that infers, merges, and evaluates user goals in real time and visualizes their progress, tested with 20 users on a writing task.

  2. SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    SEADialogues is a culturally grounded multi-turn dialogue dataset covering eight Southeast Asian languages.

  3. DialogueForge: LLM Simulation of Human-Chatbot Dialogue

    cs.CL 2025-07 conditional novelty 4.0 of 10

    DialogueForge generates synthetic human-chatbot dialogues by pitting an inquirer LLM against a responder LLM, and finds that fine-tuned small models can approach GPT-4o-level realism on LLM-judged metrics.

Pith tools