Pith. sign in

REVIEW 5 cited by

Prompt-Based Monte-Carlo Tree Search for Goal-Oriented Dialogue Policy Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13660 v2 pith:7D5MU4UC submitted 2023-05-23 cs.CL

classification cs.CL
keywords searchdialoguegoal-orientedgdp-zeromodelplanningpolicytraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress. Many approaches thus consider training neural networks to perform look-ahead search algorithms such as A* search and Monte Carlo Tree Search (MCTS). However, this training often requires abundant annotated data, which creates challenges when faced with noisy annotations or low-resource settings. We introduce GDP-Zero, an approach using Open-Loop MCTS to perform goal-oriented dialogue policy planning without any model training. GDP-Zero prompts a large language model to act as a policy prior, value function, user simulator, and system model during the tree search. We evaluate GDP-Zero on the goal-oriented task PersuasionForGood, and find that its responses are preferred over ChatGPT up to 59.32% of the time, and are rated more persuasive than ChatGPT during interactive evaluations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Dialogue agents aligned via DPO on preference pairs mined from simulated conversations improve engagement scores against the same simulator, with smaller and partially inconsistent human evaluation evidence.

  2. Tailored Conversations beyond LLMs: A RL-Based Dialogue Manager

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A hierarchical RL and meta-learning dialogue manager conditions an LLM for motivational interviewing and reports higher reward than a prompted LLM baseline in a simulated environment.

  3. Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning

    cs.CL 2025-04 conditional novelty 6.0 of 10

    UDP, a user-tailored dialogue policy planner with a diffusion-based persona portrayer and a Brownian Bridge feedback anticipator, outperforms existing planners on simulated persuasion and emotional-support tasks.

  4. Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations

    cs.CL 2026-08 reject novelty 5.0 of 10

    Preference Tree Optimization uses look-ahead simulations scored by an AI oracle to generate DPO preference data, and the resulting Motivational Interviewing agent scores higher on that same oracle than the base model.

  5. RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Adding retrieval-augmented actions and a retrieval-based factuality scorer to the rStar tree-search framework improves accuracy on six medical and commonsense QA benchmarks, reportedly matching or beating GPT-4o on se...

Pith tools