Pith. sign in

REVIEW 10 cited by

PlanGenLLMs: A Modern Survey of LLM Planning Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11221 v3 pith:PZFIAG2K submitted 2025-02-16 cs.AI cs.CL

classification cs.AIcs.CL
keywords planningcriteriallmsmakingstatesurveytasksagentic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

LLMs have immense potential for generating plans, transforming an initial world state into a desired goal state. A large body of research has explored the use of LLMs for various planning tasks, from web navigation to travel planning and database querying. However, many of these systems are tailored to specific problems, making it challenging to compare them or determine the best approach for new tasks. There is also a lack of clear and consistent evaluation criteria. Our survey aims to offer a comprehensive overview of current LLM planners to fill this gap. It builds on foundational work by Kartam and Wilkins (1990) and examines six key performance criteria: completeness, executability, optimality, representation, generalization, and efficiency. For each, we provide a thorough analysis of representative works and highlight their strengths and weaknesses. Our paper also identifies crucial future directions, making it a valuable resource for both practitioners and newcomers interested in leveraging LLM planning to support agentic workflows.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

    cs.AI 2025-10 conditional novelty 6.0 of 10

    A three-level temporal-logic safety evaluator for embodied LLM agents that checks NL-to-LTL interpretation, plan compliance, and CTL over simulated execution trees.

  2. A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models

    cs.AI 2025-09 conditional novelty 5.0 of 10

    The authors organize LLM-based time series reasoning into three exclusive topologies (direct, chain, branch) crossed with four objectives, and use them to label 125 papers, benchmarks, and resources.

  3. Towards Fully Automated Molecular Simulations: Multi-Agent Framework for Simulation Setup and Force Field Extraction

    cs.AI 2025-09 conditional novelty 5.0 of 10

    A multi-agent LLM system generates RASPA simulation inputs and extracts literature force field parameters with moderate to high accuracy on a small set of zeolite tasks.

  4. Agentic Enterprise: AI-Centric User to User-Centric AI

    cs.AI 2025-06 unverdicted novelty 5.0 of 10

    Enterprise AI should be reorganized into a user-centric market of specialized agents guided by six tenets rather than built around general-purpose assistants.

  5. Graph-Based Physics-Guided Urban PM2.5 Air Quality Imputation with Constrained Monitoring Data

    cs.LG 2025-06 conditional novelty 5.0 of 10

    GraPhy, a physics-inspired graph neural network with wind-based edge features and learnable diffusion scaling, reports the best PM2.5 imputation accuracy among six baselines on 41 sensors in Fresno, California.

  6. MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

    cs.CL 2025-07 conditional novelty 4.0 of 10

    MiroMind-M1 open-sources a two-stage SFT plus RLVR recipe with a new context-aware multi-stage policy optimization (CAMPO) that claims competitive AIME24, AIME25, and MATH500 scores among Qwen-2.5-based models.

  7. Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services

    cs.CR 2025-05 conditional novelty 4.0 of 10

    A taxonomy and auditing framework for hidden operations that opaque LLM APIs bill users for, with proposals for commitment-based, predictive, behavioral, and hardware-based verification.

  8. LLMs Between the Nodes: Community Discovery Beyond Vectors

    cs.SI 2025-07 reject novelty 3.0 of 10

    CommLLM, a two-step graph-to-text plus LLM prompting method, reports high NMI on six small networks, but its evaluation omits standard community-detection baselines and relies on a prompt tuned on one test set.

  9. Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

    cs.LG 2025-06 conditional novelty 3.0 of 10

    Under the assumption that the response reward equals the discounted sum of token rewards, response-level rewards suffice for unbiased token-level policy gradients in LLMs.

  10. Large Language Models for Planning: A Comprehensive and Systematic Survey

    cs.AI 2025-05 conditional novelty 3.0 of 10

    A structured survey of LLM planning methods, benchmarks, and interpretability work, organized around a three-way taxonomy.

Pith tools