Pith. sign in

REVIEW 2 cited by

PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User Simulator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.11534 v6 pith:GAIZK6BE submitted 2023-08-21 cs.CL cs.AI

classification cs.CLcs.AI
keywords chatgptconversationshumanmulti-rounduserbetterdialoguesgenuine
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The unparalleled performance of closed-sourced ChatGPT has sparked efforts towards its democratization, with notable strides made by leveraging real user and ChatGPT dialogues, as evidenced by Vicuna. However, due to challenges in gathering dialogues involving human participation, current endeavors like Baize and UltraChat rely on ChatGPT conducting roleplay to simulate humans based on instructions, resulting in overdependence on seeds, diminished human-likeness, limited topic diversity, and an absence of genuine multi-round conversational dynamics. To address the above issues, we propose a paradigm to simulate human behavior better and explore the benefits of incorporating more human-like questions in multi-turn conversations. Specifically, we directly target human questions extracted from genuine human-machine conversations as a learning goal and provide a novel user simulator called `Socratic'. The experimental results show our response model, `PlatoLM', achieves SoTA performance among LLaMA-based 7B models in MT-Bench. Our findings further demonstrate that our method introduces highly human-like questioning patterns and rich topic structures, which can teach the response model better than previous works in multi-round conversations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReviewInstruct: A Review-Driven Multi-Turn Conversations Generation Method for Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A review-driven multi-agent pipeline turns single-turn instruction data into harder, more diverse multi-turn dialogues and improves a Llama2-13B model on MT-Bench and MMLU-Pro.

  2. TWICE: Modeling the Temporal Evolution of Personalized User Behavior via Event-Driven Agents

    cs.IR 2025-12 reject novelty 4.0 of 10

    TWICE is an LLM framework that simulates personalized user tweets using user profiles, event-driven memory, and style rewriting; its evaluation, however, leaks the target event and lacks baselines.

Pith tools