Pith. sign in

REVIEW 3 cited by

Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.18127 v2 pith:P7CN5XXU submitted 2023-10-27 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords frameworkactionshumanlearningpoliciespromptsreasoningaction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Large language models (LLMs) demonstrate their promise in tackling complicated practical challenges by combining action-based policies with chain of thought (CoT) reasoning. Having high-quality prompts on hand, however, is vital to the framework's effectiveness. Currently, these prompts are handcrafted utilising extensive human labor, resulting in CoT policies that frequently fail to generalise. Human intervention is also required to develop grounding functions that ensure low-level controllers appropriately process CoT reasoning. In this paper, we propose a comprehensive training framework for complex task-solving, incorporating human prior knowledge into the learning of action policies. To that purpose, we offer a new leader-follower bilevel framework that is capable of learning to ask relevant questions (prompts) and subsequently undertaking reasoning to guide the learning of actions. The prompt policy is employed to make introspective revisions based on historical findings, leading the CoT process to consider the anticipated goals and generate outputs that lead to decisive, high-performing actions. The action policy subsequently learns to comprehend and integrate the CoT outputs to take actions. Our empirical data reveal that our framework outperforms leading methods in $5$ decision-making tasks such as Overcooked and FourRoom.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agentic Episodic Control

    cs.AI 2025-06 conditional novelty 6.0 of 10

    AEC couples an LLM semantic encoder, a graph working memory, and a critical-state gate to make episodic control in text-based RL more sample-efficient than standard RL baselines.

  2. IDEA: Augmenting Design Intelligence through Design Space Exploration

    cs.HC 2025-06 conditional novelty 5.0 of 10

    IDEA combines LLM-generated constraints with Monte Carlo Tree Search over a formal design space to automate design decision-making in data storytelling and pictorial visualization.

  3. HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving

    cs.RO 2025-05 conditional novelty 5.0 of 10

    The HCRMP planner feeds LLM semantic hints into state representation and critic weighting instead of letting the LLM decide actions, reporting better CARLA driving metrics.

Pith tools