Pith. sign in

REVIEW 1 cited by

Large Language Models as Commonsense Knowledge for Large-Scale Task Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14078 v2 pith:ZOVUJGWD submitted 2023-05-23 cs.RO

classification cs.RO
keywords policymodelplanningworldsearchcommonsensellmsmcts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale task planning is a major challenge. Recent work exploits large language models (LLMs) directly as a policy and shows surprisingly interesting results. This paper shows that LLMs provide a commonsense model of the world in addition to a policy that acts on it. The world model and the policy can be combined in a search algorithm, such as Monte Carlo Tree Search (MCTS), to scale up task planning. In our new LLM-MCTS algorithm, the LLM-induced world model provides a commonsense prior belief for MCTS to achieve effective reasoning; the LLM-induced policy acts as a heuristic to guide the search, vastly improving search efficiency. Experiments show that LLM-MCTS outperforms both MCTS alone and policies induced by LLMs (GPT2 and GPT3.5) by a wide margin, for complex, novel tasks. Further experiments and analyses on multiple tasks -- multiplication, multi-hop travel planning, object rearrangement -- suggest minimum description length (MDL) as a general guiding principle: if the description length of the world model is substantially smaller than that of the policy, using LLM as a world model for model-based planning is likely better than using LLM solely as a policy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling

    cs.CL 2025-07 conditional novelty 5.0 of 10

    ChiMed 2.0 is a 204.4M-character Chinese medical dataset spanning pretraining, SFT, and preference data that yields small gains on CMMLU and CEval medical subsets.

Pith tools