Pith. sign in

Paper Citation Record · LEDGER

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2506.23127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23127 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:55:06.821160Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:33:38.485506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T11:55:33.536071Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6cddb675-a768-4624-9741-83ba498cb8ce · outbound

This paper cites an unresolved cited work.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.415525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.415525Z digest=sha256:680e939bf9e0322da46c177c493bae422d1bb1e741ed0d4de6a1e7f5afe0d0e7

Observation 92f03b00-374b-4a14-9c7c-068dcff9429b · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning FireAct: Toward Language Agent Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.459541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.459541Z digest=sha256:45a3008c3ed61dbc73ffbbd6ffca3b15ab1768d0b898a1338b69a6cb9910c87a

Observation 44fdcf0d-5700-463d-831e-9d8541d497c4 · outbound

This paper cites RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.505894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.505894Z digest=sha256:0fae9963b06f6ce670e137ad97d8c226546d9636f5b5783164a58c68bc013151

Observation 46e7171b-ec46-4bee-b76d-f05f0060db78 · outbound

This paper cites Process Reward Models for LLM Agents: Practical Framework and Directions.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.569399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.569399Z digest=sha256:ccd5959d62e448c3fac872a0cb37f6634f2aed9f0df6e1bb7736c8907018253e

Observation 30bfaac9-0107-4ba4-9eac-9899aff3e25f · outbound

This paper cites DeepSeek-V3 Technical Report.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning DeepSeek-V3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.616168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.616168Z digest=sha256:fd31932e1a77ebf50f7e6cee877d03a3aa98ae4be0761223f5bb0766d69174ce

Observation 291326ee-5722-4bd0-baa2-1e87c5d81ee1 · outbound

This paper cites A survey of embodied AI: from simulators to research tasks.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning A survey of embodied AI: from simulators to research tasks

Reference 7

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T21:55:07.711625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:00.719237Z digest=sha256:0866735b750e524b5e9974cd7429e03d5f5b30fde8af79e15e3c8096bc1bec81

Observation b2be25ea-7c01-4204-a2c6-394fa2d1dabf · outbound

This paper cites an unresolved cited work.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:55:10.547481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:00.774975Z digest=sha256:bd422e1015dcf9165a0fd345833cc7896b124ee7bd89f5d1e9ca246367402055

Observation 99b878b9-8cba-4204-9b4c-e37587f6fd7a · outbound

This paper cites Chawla, Olaf Wiest, and Xiangliang Zhang.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Chawla, Olaf Wiest, and Xiangliang Zhang

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:10.308737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:00.836354Z digest=sha256:aff232afe49fccc4eab0b298aafc84cb932ada8ab10583dbdd1217152087c339

Observation 0311ce09-d251-483a-b835-3ca4da948942 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.918119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.918119Z digest=sha256:0c243ac8398c4f9f6b420e66d1ab012808eda99d0a24c8ac42482707242e6699

Observation 11797ada-e311-4259-bce7-bb4206757cc2 · outbound

This paper cites Controlling Large Language Model with Latent Actions.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Controlling Large Language Model with Latent Actions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.009132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.009132Z digest=sha256:b0c8196dac7973ae3cc93a46c702d112e3be50e48c3535cb1999c12884c375e3

Observation e10d8081-b76b-49ce-8a53-2cfd8c19ab6c · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.078098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.078098Z digest=sha256:88d0b912a71872fa3fed2aa77a95a5ee3110b26ffcea27dd76fba001c275ea5d

Observation 349398af-d142-4647-8926-ff5acfa64ff4 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.138159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.138159Z digest=sha256:693fe141a7ee90e9d4a79256bc8f7d5060572c5a902cfb24a4bdb99d4cf4784b

Observation 0f53b722-1eae-4f6b-a001-79151293c5cb · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.191538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.191538Z digest=sha256:d3aea2f103b931934e18d9a043bbc8fe49ce8ec28d881a1969f0e5f427599020

Observation 7654a8ce-2bd2-4e0f-b883-0fd401de6487 · outbound

This paper cites Let's verify step by step.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Let's verify step by step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.255445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.255445Z digest=sha256:28854275d41bc0011e3ad0507ccdf5cd9aa1225a6106ebcfbe6d5ef8d5a783c3

Observation c80cf3bd-6b7e-4a89-9abc-fc7703a8b8a9 · outbound

This paper cites QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.328590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.328590Z digest=sha256:adb303a11a8d1159c9c4594113118889c4a06356e35e0ea622945ff35f0dca58

Observation cdef8133-077c-4f9b-82d3-ff157b6b2c48 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.382412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.382412Z digest=sha256:73496ed7886ffc88b8331efe77f60d6536494d90f74d6238425de982ac7c12ca

Observation 1a76f25c-b45c-473e-ba67-e1b922a9e1ea · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.440500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.440500Z digest=sha256:64df841cbdb43b5cd0b04fd654b28cc7c8eaf29605bed2a3fcebe049b618c858

Observation aa62b6a0-3e2f-4f47-9455-ecab2d1090bc · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.518390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.518390Z digest=sha256:370b177dabe84155721ff58a94c908b02f735f60bbbff154360786352a107213

Observation 69dcfd01-d48e-4d21-85dd-be3372de2929 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.685340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.685340Z digest=sha256:1978274aa81e9ebf3a1ab5b986aa51a6e10bd3ceafc8c2e916416d3f021f106f

Observation 6f85de20-9000-4cc6-82a2-35398d56805f · outbound

This paper cites FILM: following instructions in language with modular methods.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning FILM: following instructions in language with modular methods

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:10.129092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:01.804811Z digest=sha256:5a5c1f34e413cdf296e0a32cfb4d439880e907f89ed6cc14ca0fdcd0c17c0bf0

Observation b199611a-4214-4d9e-8e4f-d1d02f70d8b6 · outbound

This paper cites Skill set optimization: Reinforcing language model behavior via transferable skills.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Skill set optimization: Reinforcing language model behavior via transferable skills

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.926905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:01.956591Z digest=sha256:8cac9a3e6fd7471df59702587158c209249daecdcd63d9815802fb3eb1eeb04d

Observation 372079f1-79b5-4900-9ebb-f38615bd7ee4 · outbound

This paper cites GPT-4 Technical Report.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:02.134672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:02.134672Z digest=sha256:048b6c1b9855ea2dc2957fc2faa8b9501cceee9d8c8cc3e6b32debe2329c481a

Observation dbde38d3-4b95-4ad0-9163-63e5af4152cf · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:02.328497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:02.328497Z digest=sha256:f248f15eefd5499baa96cdd2a1bf2a9aadfe2e659d9ffec986655095cef43f90

Observation edfb6089-6ce5-4682-bddf-b787eb2e14f1 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:02.476705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:02.476705Z digest=sha256:d7e10d5acd0bd31912d3cc092799f41e6cbcaba00036ebf71dd9e1d847e6afa1

Observation 928ca13b-4213-46a8-af18-cdbadaa4124d · outbound

This paper cites Agent planning with world knowledge model.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent planning with world knowledge model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.836841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:02.602431Z digest=sha256:8f8ad83511741f509983356ec694999bc768637b431892758937fdbc1e860c59

Observation 773b8528-02e7-4ffa-a46f-d9bbb0d3cf37 · outbound

This paper cites Tarr, William W.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Tarr, William W

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.612395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:02.766309Z digest=sha256:af4fb13157daf242046514e8c0e4f796451d1d367a482c2a9086b5f318217c68

Observation cc6260ec-5258-4a64-8d1a-6880233e3d7f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:02.882464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:02.882464Z digest=sha256:d1075bf50be638ff62c81b8d2636e115c424b0614fdae8c159b41822cf10a94b

Observation b84dae45-9bbd-4e05-a1fb-2582714e1b90 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Reflexion: language agents with verbal reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.485186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:03.022168Z digest=sha256:f9431ec15c50d28f763ae4ef8f12a9a9294acb1786980f79892f734acaa10342

Observation 88014892-07ad-4018-b6ef-a7cb5a6f0324 · outbound

This paper cites Hausknecht.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Hausknecht

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.246865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:03.145214Z digest=sha256:2bb25de151c8065b3d34ab864d68f017a55f3217810ec5294712a643acf74e03

Observation f57c87d5-c1e2-472e-a7fe-5c3815e962c6 · outbound

This paper cites Agentbank: Towards generalized LLM agents via fine-tuning on 50000+ interaction trajectories.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agentbank: Towards generalized LLM agents via fine-tuning on 50000+ interaction trajectories

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.007938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:03.254321Z digest=sha256:57dc19d67812614ba94c3528791eda66bcb41dfb1d06e28283c23bbcef23e32d

Observation 30ea569b-757d-4043-b333-015c580c87eb · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:03.345545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:03.345545Z digest=sha256:4981738ff6c56617d8d2f43f06ea09cd31c22fcb70354864aa9114c6776f0f3e

Observation 8b9603a2-5159-4b2b-a17b-79b7fc87cc1c · outbound

This paper cites Adaplanner: Adaptive planning from feedback with language models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Adaplanner: Adaptive planning from feedback with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:08.746627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:03.469279Z digest=sha256:7712c393592c119a5551976db3dfc6b8a5f61db6e1d78ba432754169f13d0839

Observation 82d41c2d-4246-44ec-bca0-328c26073f9f · outbound

This paper cites A Survey of Reasoning with Foundation Models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning A Survey of Reasoning with Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:03.633102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:03.633102Z digest=sha256:81dc33c4ae81e51001f414f1b0895344e4327bc1b09fb62e6b8d34e7676fbe7d

Observation 4237dab2-8f17-4de4-8124-79efe558c357 · outbound

This paper cites InternLM2 Technical Report.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning InternLM2 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:03.793236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:03.793236Z digest=sha256:4b1e2619b01f83dc47e0b1209d84822403b5ce6fc35906730cc37d33cfb7c3e2

Observation 70f59dfc-0d62-4f21-802e-c5b1c646d098 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:03.920658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:03.920658Z digest=sha256:7a75fa9659f1c07dc74ed88582a507ac185a3950b006b821404a72c41a77528e

Observation 5c6d6211-4907-4e1e-afef-1f97b687be2d · outbound

This paper cites STeCa: Step-level Trajectory Calibration for LLM Agent Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning STeCa: Step-level Trajectory Calibration for LLM Agent Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.117987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.117987Z digest=sha256:da1ee95fe8859a0fce8748834d1cae397b8078263b7e3eecfc5ebe94f050aeb1

Observation 3c787b02-2301-4e54-b8de-08a800880ea4 · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.280984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.280984Z digest=sha256:4f8d79995bbbe5366451711ae715115c68586687356e332a2724729ec0cf2c02

Observation 73074153-2e5e-44f9-81ac-73f5cfc3c85f · outbound

This paper cites Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.444807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.444807Z digest=sha256:23e73ac47ae1b33242f80757ebd8c45164f36e7af46fb9dc6f7d22324308ec0b

Observation 779b1c05-6328-4b05-ab3b-ac7f085a7e96 · outbound

This paper cites Jansen, Marc - Alexandre C \^ o t \' e , and Prithviraj Ammanabrolu.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Jansen, Marc - Alexandre C \^ o t \' e , and Prithviraj Ammanabrolu

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.517618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.517618Z digest=sha256:ce2c0d5e6e71c32fe34bc7598516a5f85367159c92c9dfd6883ccefd7db264bb

Observation 4f8e557f-9eeb-46fc-bf09-8e1b78caf713 · outbound

This paper cites World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.621217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.621217Z digest=sha256:a24cb483905122e21da7bbad41311fbd505121dc8e6ae4a86c1bb92e57f5cba4

Observation bb23b17a-7f74-443f-85ac-367da61f1473 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.695626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.695626Z digest=sha256:8f4eb620e90e627fc2a79d0ff621a693e6c1adceb56f7ad11ad8570ffc6179dd

Observation 36fc8b99-1a2c-46e7-ad87-53556fddc0e6 · outbound

This paper cites Chi, Quoc V.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Chi, Quoc V

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.775703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.775703Z digest=sha256:0dba11fd30518eeabbdca150ece133baab91ca51879e83b76c13fea743df0b5b

Observation 9312c6b6-0e23-4cd2-bc03-6d4bc44f2aa1 · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.850713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.850713Z digest=sha256:189a76e8f4b21064b672e70eccdaba13ab03d7cf1dc481ab6cfe5266f019c2d1

Observation f3f216b5-4b4e-4209-a147-8f13ab461ca9 · outbound

This paper cites The rise and potential of large language model based agents: a survey.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning The rise and potential of large language model based agents: a survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.969618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.969618Z digest=sha256:bc7c5d941d0a4aa88c44b30a312391a5abc099b27d0a2b166037fc58b4af89f3

Observation fbba7441-fa9f-4444-943d-2d1116594347 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.071076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.071076Z digest=sha256:726445233c8d20bbb30c328d83a4e2958c36a261639d8fe4d324aa725f43430b

Observation 9a401f42-8a84-4a96-85aa-d93346523a38 · outbound

This paper cites Watch every step! LLM agent learning via iterative step-level process refinement.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Watch every step! LLM agent learning via iterative step-level process refinement

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.154163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.154163Z digest=sha256:e9be8cf97e03f6138f4e0f29aa651fc9ccd602497e52c437007ae1a55d8013d7

Observation 61f5e7dc-8d9f-4914-9e8b-0a5c6114f23a · outbound

This paper cites Watch every step! LLM agent learning via iterative step-level process refinement.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Watch every step! LLM agent learning via iterative step-level process refinement

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:08.527617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:05.239959Z digest=sha256:16703801ff1fbd9759f6350e0e82470e32d02b8267b51f35dee2527683f53293

Observation 5ea0e976-666d-420d-a67b-5d90bd1ddfdc · outbound

This paper cites Qwen2.5 Technical Report.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.303985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.303985Z digest=sha256:584c2b22f1a4fe3b3be658fb0658ce4048cd8cae280ac25b92cb536a76bebad6

Observation 8c0eeca8-6703-4643-aacd-fdc1955a67d2 · outbound

This paper cites CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.375408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.375408Z digest=sha256:88bae8ad1e594dda264b8d17d5e13e46b010016bf5c6e9afa8ac436a0551aa20

Observation b4c4ce4d-1dc0-4c3b-9edb-93ffd23c3513 · outbound

This paper cites Narasimhan, and Yuan Cao.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Narasimhan, and Yuan Cao

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.442339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.442339Z digest=sha256:4575056c9fa79b14c5e13f6b756f1fab7be67b5000c61e7b64d43194cac59263

Observation 31beb251-bf4e-4270-a5cf-b1ac54fdd52f · outbound

This paper cites N., Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning N., Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:08.312149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:05.501822Z digest=sha256:c8e8128c0d8035c88a107ac62d524d1c94f7902b29aa026d155577ed967d86b1

Observation 3848eb40-f0f0-4ab3-af49-cebe1ec2661c · outbound

This paper cites Agent lumos: Unified and modular training for open-source language agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent lumos: Unified and modular training for open-source language agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.594415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.594415Z digest=sha256:ce54e8d6f03e8b8bd2442255e7b780fcdc7f6eb0bd8cf90113d61d9d2579cac8

Observation 8084ff0b-4e90-42fa-bc50-6cd0ad92309c · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.725140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.725140Z digest=sha256:70f72ab383ea89fce026ffdfdf197a153d7f20c373e57b66198d668258dfefe5

Observation 60ac31f0-aca5-4678-9ad2-e093e0cffafa · outbound

This paper cites Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.817610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.817610Z digest=sha256:b83eef93242b67b125dfcb00203f1c29d66bc4ad75888e5add29d25167fa4465

Observation cd366988-da8b-4428-9b9e-d5d54b3591c0 · outbound

This paper cites Agenttuning: Enabling generalized agent abilities for llms.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agenttuning: Enabling generalized agent abilities for llms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.980261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.980261Z digest=sha256:3336afe763a94a4ce005d20a5eb0d82050a065bf9d4d26e07c6cf52723c796cf

Observation ac810c3d-7994-44d2-a01e-d81a1b9d02c9 · outbound

This paper cites Enhancing decision-making for LLM agents via step-level q-value models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Enhancing decision-making for LLM agents via step-level q-value models

Reference 58

Resolution
verified exact
doi, observed 2026-08-06T21:55:07.075035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:06.053884Z digest=sha256:51276f4af98b8bfbf05ec18d9042b56a77e7826393a77e2d6489b180fc32696a

Observation cfa7755b-703c-4908-98a1-9002c37c780e · outbound

This paper cites Large language models as commonsense knowledge for large-scale task planning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Large language models as commonsense knowledge for large-scale task planning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:08.111874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:06.177700Z digest=sha256:a6d06f1770d26ab951a825bc8847fcb715674326ff48f4c65e6c7a4438ad708c

Observation c298c002-3bf0-4527-bdd5-9f85043486d1 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.305305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.305305Z digest=sha256:b86a754ca26fae9f3aa2fbc6d9d3250de854df91024ae4ac8e0cb983f73e922d

Observation 606cad4c-3865-416f-8d17-7281ef795196 · outbound

This paper cites Agents: An Open-source Framework for Autonomous Language Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agents: An Open-source Framework for Autonomous Language Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.427751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.427751Z digest=sha256:328050aa17dec869cde242df3a2978b626d2f7dbb1f25c960934f09b6d19ab12

Observation 5861ae12-b48c-443d-b637-12da6f2d9901 · outbound

This paper cites K now A gent: Knowledge-augmented planning for LLM -based agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning K now A gent: Knowledge-augmented planning for LLM -based agents

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:07.941110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T21:55:06.534526Z digest=sha256:27808e3a36374c8373868a3c332803dc3538af9e0f4857807da04a1e4779f898

Observation 61970ca3-7b6b-4135-9bf8-8e1a78af80fc · outbound

This paper cites write newline.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning write newline

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.577975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.577975Z digest=sha256:1a4c0f182184f00ee37235ae2a6f11dfef61b01e5460a05ac9dad53e063a5e26

Observation bb9523f3-12e8-4d73-a115-2a0433ca4c7a · outbound

This paper cites @esa (Ref.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning @esa (Ref

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.675280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.675280Z digest=sha256:5459ff9556d373f207c9a893481e73222978214164b81b9ca9b22ed97eb0486e

Observation 048ecf1a-75c5-496b-94d8-7d7281645b8f · outbound

This paper cites an unresolved cited work.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.732398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.732398Z digest=sha256:bb3fba06371273c0de790d87a57170979e2b8c517da1b9d83906549f6b4d43dd

Observation 6f087a48-75c0-4512-99f8-8e14c8b6bd0f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.821160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.821160Z digest=sha256:d1cbe5c0356e2e1f737b69158ed4118c1e1acd3d549d56a50ebc7e91f434d703

Pith citing papers

Observation 6f8873d8-e2c1-4d93-b663-d1c42fe201b3 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:38.485506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:38.485506Z digest=sha256:086b1c8f28ec6b5588838867a7b1f45691e121b023f27138721f7bd1451dfa87

Observation d509ff7e-1e52-432c-8d78-a450afbf3bec · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.429654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:8d0289a82f9609f64f1420fed52e0c953723825dcbe24c81a85c095cfd6649bd

Observation 6ce41588-75c6-48de-8f13-a86b0b9fbce5 · inbound

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning cites this paper.

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:55:33.538357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T11:52:58.630590Z digest=sha256:ea0398802bd17bc30cd54b9d5c8c92a0740a80b14efb223ddc13974d144063ed