REVIEW 5 cited by
Prompt-Based Monte-Carlo Tree Search for Goal-Oriented Dialogue Policy Planning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress. Many approaches thus consider training neural networks to perform look-ahead search algorithms such as A* search and Monte Carlo Tree Search (MCTS). However, this training often requires abundant annotated data, which creates challenges when faced with noisy annotations or low-resource settings. We introduce GDP-Zero, an approach using Open-Loop MCTS to perform goal-oriented dialogue policy planning without any model training. GDP-Zero prompts a large language model to act as a policy prior, value function, user simulator, and system model during the tree search. We evaluate GDP-Zero on the goal-oriented task PersuasionForGood, and find that its responses are preferred over ChatGPT up to 59.32% of the time, and are rated more persuasive than ChatGPT during interactive evaluations.
Forward citations
Cited by 5 Pith papers
-
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
Dialogue agents aligned via DPO on preference pairs mined from simulated conversations improve engagement scores against the same simulator, with smaller and partially inconsistent human evaluation evidence.
-
Tailored Conversations beyond LLMs: A RL-Based Dialogue Manager
A hierarchical RL and meta-learning dialogue manager conditions an LLM for motivational interviewing and reports higher reward than a prompted LLM baseline in a simulated environment.
-
Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning
UDP, a user-tailored dialogue policy planner with a diffusion-based persona portrayer and a Brownian Bridge feedback anticipator, outperforms existing planners on simulated persuasion and emotional-support tasks.
-
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
Preference Tree Optimization uses look-ahead simulations scored by an AI oracle to generate DPO preference data, and the resulting Motivational Interviewing agent scores higher on that same oracle than the base model.
-
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models
Adding retrieval-augmented actions and a retrieval-based factuality scorer to the rStar tree-search framework improves accuracy on six medical and commonsense QA benchmarks, reportedly matching or beating GPT-4o on se...
Discussion (0). Continue with ORCID to comment.