Pith. sign in

Paper Citation Record · LEDGER

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study

As of 19 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2507.23589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23589 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:45:59.609092Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:40:52.854907Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy14
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 49bc82a5-2191-4a10-96b6-60a483259a01 · outbound

This paper cites The fast downward planning system.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study The fast downward planning system

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.456686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.456686Z digest=sha256:f19eb663078a5cb4aed9a89a7f731cecdcf6a3be11ffe742d63697163bd12144

Observation 6250c9d6-e2cf-44e2-bfec-06ef27bb4fde · outbound

This paper cites PDDL —the planning domain definition language.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study PDDL —the planning domain definition language

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.495048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.461111Z digest=sha256:2df65806b242e3ff846ac4abbd6ea93e820988b77516fd26daf7578202567b1e

Observation 18b393af-a807-4e95-bc45-bdf63897c314 · outbound

This paper cites Chi, Quoc V.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Chi, Quoc V

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.484107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.466183Z digest=sha256:4b38ead45ef3010fdbd51bd06fa949cce9ecde3ce938d9a3b622b4ce0c88ddb2

Observation 0a588ba2-55db-406c-a35a-05889c642f88 · outbound

This paper cites ReAct : Synergizing reasoning and acting in language models.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study ReAct : Synergizing reasoning and acting in language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.472237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.473289Z digest=sha256:2fa229a3091c99c19bb97040c83037223d96a045e27aa8f5ff8c6048cc5d7887

Observation 5cce2666-b6d6-45a4-92a8-9526e956dbba · outbound

This paper cites Sadler, Wei-Lun Chao, and Yu Su.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Sadler, Wei-Lun Chao, and Yu Su

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-06T10:46:00.287389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.477684Z digest=sha256:14343aef10b51c2f334fc2e7f00a0c1544690566d44dbf14b439638a3e9d73a9

Observation d60c3fc9-d942-4033-9d13-bf5268cf55da · outbound

This paper cites Generating Executable Action Plans with Environmentally-Aware Language Models.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Generating Executable Action Plans with Environmentally-Aware Language Models

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-06T10:46:00.199508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.481762Z digest=sha256:0f3642601dd80346aaf31979af945f3e50637b9368d8e2970d8e53cde845433c

Observation 042a9429-61c7-4afc-b456-95c0e358e9d6 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.485930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.485930Z digest=sha256:24aca2aea9612987cd109cc6f3fa069703f035f8b51d5aa213ba994b6b0989cc

Observation b94b3c2c-3e92-4ff1-931a-a6c8ed2905c0 · outbound

This paper cites an unresolved cited work.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T10:46:00.460490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.490289Z digest=sha256:1bff2ab419582aff633c0371af1c8d6e4491058c51b63283909a0bef016d5d24

Observation 496886cf-8f00-474e-b00a-537900bc8640 · outbound

This paper cites Can Large Language Models Reason and Plan? Annals of the New York Academy of Sciences, 2024.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Can Large Language Models Reason and Plan? Annals of the New York Academy of Sciences, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.493963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.493963Z digest=sha256:64072b228cae16f5f7c259b8ddf628749e53983adfca9f6a44f6c207c26c120d

Observation b602f388-51a7-4dc8-812b-22546b0d9c52 · outbound

This paper cites Leveraging environment interaction for automated pddl translation and planning with large language models.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Leveraging environment interaction for automated pddl translation and planning with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.448753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.497857Z digest=sha256:2d675229b780d7f288ab5702da48941b27cd1e470b6b205b0657e6e1b9bea103

Observation 1fce17a5-63fd-4e91-8cbe-5a1404e32ec2 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.501390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.501390Z digest=sha256:7d9ff574849294930174b611e338959fc8c6c571514f83ba1f648da2b799fdd0

Observation 13e42e2b-9e24-4c0d-9ffb-ea659bf1feff · outbound

This paper cites Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.436105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.505767Z digest=sha256:65fb2f0d0368c0cf4e962610d875bab66752fabfe5646abd764d8c5644f58ec2

Observation 0b91cd9b-ab87-4a95-b070-b8200e55736f · outbound

This paper cites Can We Rely on LLM Agents to Draft Long-Horizon Plans? Let's Take TravelPlanner as an Example.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Can We Rely on LLM Agents to Draft Long-Horizon Plans? Let's Take TravelPlanner as an Example

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.509242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.509242Z digest=sha256:747d437a0c3a9dccf79d4e10e4c931fa54a1a5a9e70b04c6486af108f5742811

Observation e48dc1af-e625-4013-8d95-615a6cbd5cf3 · outbound

This paper cites an unresolved cited work.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T10:46:00.424534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.513203Z digest=sha256:e3e33149a26150a9976f19dee8fc9a859ffc238c35eb00828bab935f383f85e3

Observation 36501121-d7b5-4b19-89eb-f6865d6eeabe · outbound

This paper cites AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.516565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.516565Z digest=sha256:0c231762d5684ced7a15b4cad39dc3c8181fe513e7a77d013bb8a49be6fd09e5

Observation bab11192-388a-4bf8-a8e3-c71eea454a49 · outbound

This paper cites On the planning abilities of large language models: A critical investigation.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study On the planning abilities of large language models: A critical investigation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.412846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.520497Z digest=sha256:4396e6f207eea9e91fa2519b03a1f361aec920fa89b1261f02c0e250d96e9b06

Observation 8274e362-d01f-461d-a016-4a92c8a5a5a0 · outbound

This paper cites A Framework for Neurosymbolic Robot Action Planning using Large Language Models.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study A Framework for Neurosymbolic Robot Action Planning using Large Language Models

Reference 17

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T10:46:00.049165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.524162Z digest=sha256:29460695c4e27dc7335fa06571368a66348732d65c07eb3d627b06c487be37c1

Observation db6a0d60-8e6a-4c98-87dd-f1dd64a146f9 · outbound

This paper cites Automating the Generation of Prompts for LLM-based Action Choice in PDDL Planning.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Automating the Generation of Prompts for LLM-based Action Choice in PDDL Planning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.527985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.527985Z digest=sha256:d460cfb4a3372b35322fdfe5bd9f55c4f74fae0257dc7dcd484851a567980cd3

Observation d8388aa4-0d24-4a7e-9735-fc68e4aefe0f · outbound

This paper cites Tenenbaum, Leslie Pack Kaelbling, and Michael Katz.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Tenenbaum, Leslie Pack Kaelbling, and Michael Katz

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.531672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.531672Z digest=sha256:0cde10add0912f5b052e42fdd7851d9b984455d203f2cdfded16b7aec6f5c8ec

Observation 5b6846e1-d56b-4d9f-9728-f48209d0a078 · outbound

This paper cites Fast and Accurate Task Planning using Neuro-Symbolic Language Models and Multi-level Goal Decomposition.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Fast and Accurate Task Planning using Neuro-Symbolic Language Models and Multi-level Goal Decomposition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.539657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.539657Z digest=sha256:cf88ebfe6703ee563c79861cfb50ddfff29845efd3d5a5c9f1f73180e6b3fe4f

Observation 5670468d-2670-4b04-942b-61f796e6e533 · outbound

This paper cites CoPAL: Corrective Planning of Robot Actions with Large Language Models.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study CoPAL: Corrective Planning of Robot Actions with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.543728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.543728Z digest=sha256:d3a853f0ea48d8872b3cb2fc7c0f4842288b38dafeca71f545a7cc753379a112

Observation 363bcc86-bc10-4b74-9ece-199e8ab32e95 · outbound

This paper cites Saycanpay: Heuristic planning with large language models using learnable domain knowledge.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Saycanpay: Heuristic planning with large language models using learnable domain knowledge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.547535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.547535Z digest=sha256:f81dec8123538adf8b724b99108961093e200d991f89e3c02fb763ede0c2385b

Observation 05e5e0b6-9f30-40f4-a66f-67e5ab375afd · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study PaLM-E: An Embodied Multimodal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.551279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.551279Z digest=sha256:50222f77fd641565cccacde473e83272597718666a410c2450a8fb42f4cb047e

Observation ba4a841c-c3ee-43a7-8feb-4f039cf70fa1 · outbound

This paper cites PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.399627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.555147Z digest=sha256:4451620c3ebe1c0deb4c57210a36e96632ddbd721219281a08da6472b3adc6b5

Observation fcc2a004-06f6-47e3-a5c9-d9af1404e752 · outbound

This paper cites LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.558754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.558754Z digest=sha256:c16da25602a0f94b875b955b129ef788445b6df99b459fa45d793bafcf9577c7

Observation 5572bc3b-ed74-41cd-b8c6-7997792d9df3 · outbound

This paper cites Littman, and Stephen H.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Littman, and Stephen H

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.562768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.562768Z digest=sha256:19e100e5fd071763187a67d8b4023a80d4026612cf667018d1f0493411643ee6

Observation 249cde25-c9bc-4384-9b69-aecd0d6ffe01 · outbound

This paper cites NL2Plan: Robust LLM-Driven Planning from Minimal Text Descriptions.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study NL2Plan: Robust LLM-Driven Planning from Minimal Text Descriptions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.566806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.566806Z digest=sha256:55affbdbca674769d0d29ac96685269ebed74e663d6e2442749b60e4e8a36ed8

Observation c995e071-1e8c-4969-9b34-55c5c68ada16 · outbound

This paper cites NATURAL PLAN: Benchmarking LLMs on Natural Language Planning.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study NATURAL PLAN: Benchmarking LLMs on Natural Language Planning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.570552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.570552Z digest=sha256:7fbf80ea8291742af6707ce893c1e80a29030a11a20c8c83d904120416869c86

Observation 2f58b833-7ab4-493a-b090-abb03732f5e3 · outbound

This paper cites Open Grounded Planning: Challenges and Benchmark Construction.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Open Grounded Planning: Challenges and Benchmark Construction

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:45:59.692535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.574661Z digest=sha256:45e056bc1fd8825b7ac65f1b4bfbe134aafebe9523625c979aede189fd1c7854

Observation 49ed1861-dc8d-48b4-b69d-011a47b16ba5 · outbound

This paper cites CaT-Bench: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study CaT-Bench: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.578372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.578372Z digest=sha256:f9bb25ad60cda4907b49cd7d7b19c4887536c6ab2c811bc038a2901cf12b1a5c

Observation f79db19d-e568-42e8-8d85-f03158156c23 · outbound

This paper cites The production ai platform built for developers, 2024.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study The production ai platform built for developers, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.385229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.582074Z digest=sha256:e265077ebbe014f116c57b42e876e348a01a7e8338aa4e65d4449a1de6ad237f

Observation 3dbe9cab-60d4-4cea-9734-ffafcf100a96 · outbound

This paper cites an unresolved cited work.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.585767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.585767Z digest=sha256:0fcb87d7cdd0e48761e9320e9106622f74edfd853f9aa47331a172d19dc3b783

Observation f70cca8e-7b30-4500-a07a-208a8e730f15 · outbound

This paper cites Deepseek vs.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Deepseek vs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.365312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.589471Z digest=sha256:cdda69077cdd00f88f05fc6b54f51772c388f45926c174d42b4d4c1e72f8760d

Observation b8e668ac-58cc-4251-ae55-91003419891c · outbound

This paper cites Meta releases new llama 3.1 models, including highly anticipated 405b parameter variant.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Meta releases new llama 3.1 models, including highly anticipated 405b parameter variant

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.354585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.592960Z digest=sha256:e6e854b025330da2728f1dc7ef91af2c4aad196d86a266aab3c11b4dda44a9c8

Observation 5b54945e-3975-4b79-a3d1-174f57b3eda5 · outbound

This paper cites Gemini 2.0 flash thinking experimental: A guide with examples.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Gemini 2.0 flash thinking experimental: A guide with examples

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.339061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.597071Z digest=sha256:5f5f762634a142b32d12c788f8ce6943657a1fa43d3c16420b09e4170e71fd16

Observation c3c1fb54-4c33-48dc-909b-1ba87abc77d6 · outbound

This paper cites Comparing claude 3.7 sonnet, claude 3.5 sonnet, openai o3-mini, deepseek r1, and grok 3 beta.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Comparing claude 3.7 sonnet, claude 3.5 sonnet, openai o3-mini, deepseek r1, and grok 3 beta

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.326857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.600996Z digest=sha256:770132f5a5ababf4f4671c532cdf524b5b6f38a29ce062a5ca1da612b685d252

Observation a5ef1c3a-d9fa-467d-ae83-64385cf1b4a7 · outbound

This paper cites Grok 3 beta—the age of reasoning agents.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Grok 3 beta—the age of reasoning agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.313588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.605454Z digest=sha256:2099bb1644b411b954cea127ad9881c393c396262f8fcd03ee62edba515e2da9

Observation 74ab2bf0-21c3-4379-856f-31c298bd3518 · outbound

This paper cites Claude 3.7 sonnet and claude code.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study Claude 3.7 sonnet and claude code

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:46:00.300041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:45:59.609092Z digest=sha256:b174e52cd5c207782a0c35cf3adbe926c36c0ad173bbdfd1ce6c141d649a01e5

Pith citing papers

Observation e657ef9d-ead4-462d-8171-d0d281ced4b2 · inbound

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards cites this paper.

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:45:21.129130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T04:40:52.854907Z digest=sha256:c188a2ce471b528d0507c01c892111b867a2edd423400a8c68b3ad8094240aa6