Pith. sign in

REVIEW 15 cited by

A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.18013 v2 pith:7ERAR5JA submitted 2024-02-28 cs.CL cs.AI

classification cs.CLcs.AI
keywords dialoguesystemsmulti-turnllmsrecentadvancesllm-basedresearch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This survey provides a comprehensive review of research on multi-turn dialogue systems, with a particular focus on multi-turn dialogue systems based on large language models (LLMs). This paper aims to (a) give a summary of existing LLMs and approaches for adapting LLMs to downstream tasks; (b) elaborate recent advances in multi-turn dialogue systems, covering both LLM-based open-domain dialogue (ODD) and task-oriented dialogue (TOD) systems, along with datasets and evaluation metrics; (c) discuss some future emphasis and recent research problems arising from the development of LLMs and the increasing demands on multi-turn dialogue systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems

    cs.MA 2026-02 conditional novelty 6.0 of 10

    A perturbation-based framework measures how value opinions propagate through multi-agent LLM systems, revealing that susceptibility varies by value, model, and topology.

  2. Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection

    cs.CL 2026-02 unverdicted novelty 6.0 of 10

    Token Sparse Attention uses dynamic per-head token compression and decompression during attention to achieve up to 3.23x speedup at 128K context with under 1% accuracy loss.

  3. Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study

    cs.AI 2025-06 conditional novelty 6.0 of 10

    ChatGPT 4o and o1-mini chose more risk-averse lottery options than real respondents in Sydney, Hong Kong, Dhaka, and Nanjing; o1-mini was closer to humans, and Chinese prompts widened the gap.

  4. Proactive Guidance of Multi-Turn Conversation in Industrial Search

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A framework combining goal-adaptive supervised fine-tuning and click-based reinforcement learning improves proactive guidance quality and speed in an industrial search assistant.

  5. Evaluating the Sensitivity of LLMs to Prior Context

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Prior conversational context, especially from a different knowledge domain, can sharply reduce LLM multiple-choice accuracy, and repeating the task near the query mitigates the drop.

  6. Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems

    cs.MA 2025-05 conditional novelty 6.0 of 10

    Moderately sparse communication topologies balance error suppression and insight propagation in LLM multi-agent systems, and the proposed EIB-Learner learns such topologies to outperform prior methods.

  7. Cracking Aegis: An Adversarial LLM-based Game for Raising Awareness of Vulnerabilities in Privacy Protection

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Cracking Aegis, an adversarial LLM-driven dialogue game, led players to use manipulative language strategies and to self-report stronger awareness of privacy vulnerabilities after a single session.

  8. A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PERSONACONVBENCH is a new Reddit-based benchmark showing that LLMs predict sentiment, community scores, and next replies better when given a user's multi-turn conversation history, and it releases public data and code.

  9. Implicit Reasoning in Large Language Models: A Comprehensive Survey

    cs.CL 2025-09 conditional novelty 5.0 of 10

    A survey organizing implicit (silent) reasoning in LLMs into three execution paradigms, plus evidence, benchmarks, and challenges.

  10. Separation Logic of Generic Resources via Sheafeology

    cs.LO 2025-08 unverdicted novelty 5.0 of 10

    Sheafeology uses sheaf categories to make first-order logic resource-aware, yielding separation logics for generic resources such as memory and random variables.

  11. Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents

    cs.CL 2025-07 conditional novelty 5.0 of 10

    H-MEM organizes LLM agent memory into a four-level semantic hierarchy with pointer-based coarse-to-fine retrieval, improving average LoCoMo QA scores over five baselines while cutting retrieval cost.

  12. A Novel Self-Evolution Framework for Large Language Models

    cs.CL 2025-07 reject novelty 4.0 of 10

    A dual-phase framework that uses a Censor satisfaction scorer to expand training data and then applies SFT plus frequency-weighted DPO, reporting benchmark gains over SFT, PO, and memory baselines.

  13. An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite Individuals

    cs.CL 2025-06 conditional novelty 4.0 of 10

    An evolutionary reinforcement learning method with elite individual injection improves task-oriented dialogue policy performance on four datasets.

  14. Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery

    cs.AI 2025-07 reject novelty 3.0 of 10

    A position paper with a 33-patient pilot argues LLM phone agents can make routine monitoring cheaper, but the savings are assumed rather than measured.

  15. Why do AI agents communicate in human language?

    cs.AI 2025-06 conditional novelty 3.0 of 10

    The paper argues that natural language is structurally mismatched to LLM internal representations, so future AI agents should abandon it for inter-agent communication and train models with structured communication primitives.

Pith tools