REVIEW 6 cited by
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs), optimized through human feedback, have rapidly emerged as a leading paradigm for developing intelligent conversational assistants. However, despite their strong performance across many benchmarks, LLM-based agents might still lack conversational skills such as disambiguation -- when they are faced with ambiguity, they often overhedge or implicitly guess users' true intents rather than asking clarification questions. Under task-specific settings, high-quality conversation samples are often limited, constituting a bottleneck for LLMs' ability to learn optimal dialogue action policies. We propose Action-Based Contrastive Self-Training (ACT), a quasi-online preference optimization algorithm based on Direct Preference Optimization (DPO), that enables data-efficient dialogue policy learning in multi-turn conversation modeling. We demonstrate ACT's efficacy under in data-efficient tuning scenarios, even when there is no action label available, using multiple real-world conversational tasks: tabular-grounded question-answering, machine reading comprehension, and AmbigSQL, a novel task for disambiguating information-seeking requests for complex SQL generation towards data analysis agents. Additionally, we propose evaluating LLMs' ability to function as conversational agents by examining whether they can implicitly recognize and reason about ambiguity in conversation. ACT demonstrates substantial conversation modeling improvements over standard tuning approaches like supervised fine-tuning and DPO.
Forward citations
Cited by 6 Pith papers
-
Bridging Compute- and Data-Optimal Pretraining
Pretraining loss obeys a single law in which repeated or paraphrased tokens count as η(N, data-per-parameter, expansion-ratio) fresh tokens, with total effective data saturating as derived tokens grow.
-
Teaching Language Models To Gather Information Proactively
Rewarding questions for eliciting genuinely new information trains a small model to outperform larger models at proactive clarification and downstream writing quality.
-
CollabLLM: From Passive Responders to Active Collaborators
CollabLLM computes multiturn-aware rewards by forward-simulating future user turns, then fine-tunes the LLM with RL to ask clarifying questions and guide users to their goals.
-
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
Adding transcription and answer-selection auxiliary tasks to a fixed spoken-QA dataset improves MLLM performance with less data, validated on three datasets and a new ASK-QA benchmark.
-
Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms
A survey that proposes the Generalist Virtual Agent concept and taxonomies for agent environments, tasks, perceptions, actions, models, and evaluation, concluding that real-world-like environments favor human-like int...
-
A Survey on Multi-Turn Interaction Capabilities of Large Language Models
A comprehensive review of how large language models are evaluated, trained, and improved for multi-turn interaction, organized into evaluation practices, core capabilities, and general algorithms.
Discussion (0). Continue with ORCID to comment.