Pith. sign in

Paper Citation Record · LEDGER

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2506.23127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23127 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:55:06.821160Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:33:38.485506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T11:55:33.536071Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6cddb675-a768-4624-9741-83ba498cb8ce · outbound

This paper cites an unresolved cited work.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.415525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.415525Z digest=sha256:17b8a09f04f4d10ae286fab67cdef13cc11f253921df021d87f8c79d7d710cdb

Observation 92f03b00-374b-4a14-9c7c-068dcff9429b · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning FireAct: Toward Language Agent Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.459541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.459541Z digest=sha256:4a4f3eb7701bdf7d589f3af27d092f8ba91e993d8aef3dbed9d7ba8ff6106d1c

Observation 44fdcf0d-5700-463d-831e-9d8541d497c4 · outbound

This paper cites RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.505894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.505894Z digest=sha256:23e3f889be910900c0857a8f4cdce36c3caebfadf37eed537aae3c8d787b46c4

Observation 46e7171b-ec46-4bee-b76d-f05f0060db78 · outbound

This paper cites Process Reward Models for LLM Agents: Practical Framework and Directions.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.569399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.569399Z digest=sha256:e79f9eb2a20e9e0221b4bcd1947fe9f0f838249138b727d74acc02e8722951be

Observation 30bfaac9-0107-4ba4-9eac-9899aff3e25f · outbound

This paper cites DeepSeek-V3 Technical Report.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning DeepSeek-V3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.616168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.616168Z digest=sha256:3edfc2ad2c882666d5356530ca97f938657259424eb8d41e7920f0706b308945

Observation 291326ee-5722-4bd0-baa2-1e87c5d81ee1 · outbound

This paper cites A survey of embodied AI: from simulators to research tasks.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning A survey of embodied AI: from simulators to research tasks

Reference 7

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T21:55:07.711625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:00.719237Z digest=sha256:206a8547010d0264d1331a07437cc45c02e7f98851450b65e0e8a031965b021d

Observation b2be25ea-7c01-4204-a2c6-394fa2d1dabf · outbound

This paper cites an unresolved cited work.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:55:10.547481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:00.774975Z digest=sha256:38b5af03f3afa3028f6dd5fad9329318cb15202aef301926e557e3a58b0e51e7

Observation 99b878b9-8cba-4204-9b4c-e37587f6fd7a · outbound

This paper cites Chawla, Olaf Wiest, and Xiangliang Zhang.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Chawla, Olaf Wiest, and Xiangliang Zhang

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:10.308737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:00.836354Z digest=sha256:cf7dcf8c24de8f126700bf45db6826bdbaaf364c839d9a4b9779aeb772ff4b7a

Observation 0311ce09-d251-483a-b835-3ca4da948942 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:00.918119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:00.918119Z digest=sha256:55adb39f35f0af1c871747a2a299e636e9d9dc0ffcdabd2870627462c8276ede

Observation 11797ada-e311-4259-bce7-bb4206757cc2 · outbound

This paper cites Controlling Large Language Model with Latent Actions.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Controlling Large Language Model with Latent Actions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.009132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.009132Z digest=sha256:cb3917677e171e028b12d49256db999b35e04e062ed65791a98003d56017d24e

Observation e10d8081-b76b-49ce-8a53-2cfd8c19ab6c · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.078098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.078098Z digest=sha256:013130111eb59c1f6f987d4a844414f354f2de14977515658a7bc145a17c5fd4

Observation 349398af-d142-4647-8926-ff5acfa64ff4 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.138159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.138159Z digest=sha256:8ddde05e9b0d76029500995b9742003e926353ccebcdc454532359bbf1d5895c

Observation 0f53b722-1eae-4f6b-a001-79151293c5cb · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.191538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.191538Z digest=sha256:5a677750b5fb6903089cd8af49caa304d0da78f97c9046a466e31874c8fc412c

Observation 7654a8ce-2bd2-4e0f-b883-0fd401de6487 · outbound

This paper cites Let's verify step by step.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Let's verify step by step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.255445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.255445Z digest=sha256:370ecd7dfb6de707bcd61eb961a2e92300a20aadae04d8606c82574e4f53abea

Observation c80cf3bd-6b7e-4a89-9abc-fc7703a8b8a9 · outbound

This paper cites QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.328590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.328590Z digest=sha256:2b9a2d53645e7d673391d7270b53ac4a7d626a85b22867c5dd3085eb21eeaa41

Observation cdef8133-077c-4f9b-82d3-ff157b6b2c48 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.382412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.382412Z digest=sha256:d4ab57fe8e83a475b806633cd4e791bf445ef1afe85b4f7c91c6ae0b3f2cada6

Observation 1a76f25c-b45c-473e-ba67-e1b922a9e1ea · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.440500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.440500Z digest=sha256:4657577ce7b5126f116d1300407279d37c51902a61922b5a467a13b039594b5e

Observation aa62b6a0-3e2f-4f47-9455-ecab2d1090bc · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.518390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.518390Z digest=sha256:b318c0b426020368d1973578a92aec563e935c0eb8b299df36fcf5ef9f969fd7

Observation 69dcfd01-d48e-4d21-85dd-be3372de2929 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.685340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.685340Z digest=sha256:f70a09f2beebad8d12a02dada77456e3a2246614d9996cdbd246d29f38c1ff23

Observation 6f85de20-9000-4cc6-82a2-35398d56805f · outbound

This paper cites FILM: following instructions in language with modular methods.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning FILM: following instructions in language with modular methods

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:10.129092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:01.804811Z digest=sha256:f9d90547edfb0a7f92c04b29486368e22cd5d4e0c39ff7f710e5338622830973

Observation b199611a-4214-4d9e-8e4f-d1d02f70d8b6 · outbound

This paper cites Skill set optimization: Reinforcing language model behavior via transferable skills.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Skill set optimization: Reinforcing language model behavior via transferable skills

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.926905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:01.956591Z digest=sha256:3c42c9404d0fb90b26c9483f63b07c891cbcba81d649a483d0781545fa484dd1

Observation 372079f1-79b5-4900-9ebb-f38615bd7ee4 · outbound

This paper cites GPT-4 Technical Report.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:02.134672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:02.134672Z digest=sha256:39e893a2ac9838f9ec30bf63e25b817c737e8e2f179029c763c9da71181cc438

Observation dbde38d3-4b95-4ad0-9163-63e5af4152cf · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:02.328497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:02.328497Z digest=sha256:910aab9a828c817eb54c071dfa71d305f9013f77f453296242ff2bf5940687d1

Observation edfb6089-6ce5-4682-bddf-b787eb2e14f1 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:02.476705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:02.476705Z digest=sha256:af91ec92187eb41315be40794569a3b605ae361ebd8876cb6b650e0adf8114ed

Observation 928ca13b-4213-46a8-af18-cdbadaa4124d · outbound

This paper cites Agent planning with world knowledge model.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent planning with world knowledge model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.836841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:02.602431Z digest=sha256:3d2e2c38613dae3449e486520c38e95ac19f6b17e9a969a9522b91a6e9dbddc1

Observation 773b8528-02e7-4ffa-a46f-d9bbb0d3cf37 · outbound

This paper cites Tarr, William W.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Tarr, William W

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.612395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:02.766309Z digest=sha256:9c576e4ba78759d8af950252eeaa2c7563f6aaf9f3e6085f84b2afd23de78cdb

Observation cc6260ec-5258-4a64-8d1a-6880233e3d7f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:02.882464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:02.882464Z digest=sha256:b0cc9f5eebce419145016f8a47ebb0abbb591992a5b1f9da033afd053df072d8

Observation b84dae45-9bbd-4e05-a1fb-2582714e1b90 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Reflexion: language agents with verbal reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.485186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:03.022168Z digest=sha256:7c6ad52d54927cf260db91a04d3f39417aa32cee7bc288876a81a832cd08a022

Observation 88014892-07ad-4018-b6ef-a7cb5a6f0324 · outbound

This paper cites Hausknecht.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Hausknecht

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.246865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:03.145214Z digest=sha256:23e24900a6d357400c6c2fc5aad60b9bc8a33ccbb845bcf70efc2e7df7fa9e28

Observation f57c87d5-c1e2-472e-a7fe-5c3815e962c6 · outbound

This paper cites Agentbank: Towards generalized LLM agents via fine-tuning on 50000+ interaction trajectories.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agentbank: Towards generalized LLM agents via fine-tuning on 50000+ interaction trajectories

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:09.007938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:03.254321Z digest=sha256:a1ac48f97d3719693aab04cd42e1f74363675fcc295d1ad752c3af6b2403b149

Observation 30ea569b-757d-4043-b333-015c580c87eb · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:03.345545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:03.345545Z digest=sha256:cc4d75eb6cfa361facf6b4f06fc11720b9f0a0fc0c95f5d04044652c6c65e68f

Observation 8b9603a2-5159-4b2b-a17b-79b7fc87cc1c · outbound

This paper cites Adaplanner: Adaptive planning from feedback with language models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Adaplanner: Adaptive planning from feedback with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:08.746627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:03.469279Z digest=sha256:e0f74c4ffedccf4df725245f45e0dd0d7239dcc8545ddd67f97fa7eb4b18dba3

Observation 82d41c2d-4246-44ec-bca0-328c26073f9f · outbound

This paper cites A Survey of Reasoning with Foundation Models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning A Survey of Reasoning with Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:03.633102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:03.633102Z digest=sha256:9ccdd0ec4808bf303558f1e48771b6a3751cb4d00fe48c3a2837c2af898106cd

Observation 4237dab2-8f17-4de4-8124-79efe558c357 · outbound

This paper cites InternLM2 Technical Report.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning InternLM2 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:03.793236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:03.793236Z digest=sha256:f3795d39b21815d4ed951221cb51af332ae4b9336a685b34a3e1b6eac38e8c28

Observation 70f59dfc-0d62-4f21-802e-c5b1c646d098 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:03.920658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:03.920658Z digest=sha256:99ffc9edf99442bc45c8019f10cbd1f249d68dd7690fe788d3311075c22d2d22

Observation 5c6d6211-4907-4e1e-afef-1f97b687be2d · outbound

This paper cites STeCa: Step-level Trajectory Calibration for LLM Agent Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning STeCa: Step-level Trajectory Calibration for LLM Agent Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.117987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.117987Z digest=sha256:663aac2664f11e70659d34545e803d4657788cb207024c146a71218fbae83553

Observation 3c787b02-2301-4e54-b8de-08a800880ea4 · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.280984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.280984Z digest=sha256:6a77662adab1bafa73ee2cceca82108ec7b9b12523df7edf44eb1c043c01e822

Observation 73074153-2e5e-44f9-81ac-73f5cfc3c85f · outbound

This paper cites Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.444807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.444807Z digest=sha256:de7417e3e3409f6387f8d999a3a29ed901a89945d56e7b10f68efad44969b9af

Observation 779b1c05-6328-4b05-ab3b-ac7f085a7e96 · outbound

This paper cites Jansen, Marc - Alexandre C \^ o t \' e , and Prithviraj Ammanabrolu.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Jansen, Marc - Alexandre C \^ o t \' e , and Prithviraj Ammanabrolu

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.517618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.517618Z digest=sha256:048e1f40d6c4b21bfdbfd9f31e6f9735b7761c9a2b844225d1b4750390a05f53

Observation 4f8e557f-9eeb-46fc-bf09-8e1b78caf713 · outbound

This paper cites World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.621217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.621217Z digest=sha256:725f7facd8b53cac20217dc9c8022cfe96970186dff764c81f80b715f7c9be5b

Observation bb23b17a-7f74-443f-85ac-367da61f1473 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.695626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.695626Z digest=sha256:896f5bbb5b6d0395b0ec7a937b8aea832fd8f27c2b2279709b2c929cd4090b3e

Observation 36fc8b99-1a2c-46e7-ad87-53556fddc0e6 · outbound

This paper cites Chi, Quoc V.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Chi, Quoc V

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.775703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.775703Z digest=sha256:3218dd8dee32315fc3c5342080b5bc13f659fe94c8842c5b6e3ce944f924252d

Observation 9312c6b6-0e23-4cd2-bc03-6d4bc44f2aa1 · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.850713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.850713Z digest=sha256:36eada45c0fa7e1efa90210565f54852a0e2b3ee7e10ba7468115f4fbbe92eb2

Observation f3f216b5-4b4e-4209-a147-8f13ab461ca9 · outbound

This paper cites The rise and potential of large language model based agents: a survey.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning The rise and potential of large language model based agents: a survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.969618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.969618Z digest=sha256:1e0c15c4d855bb0055c474053eddceac496e5813a10b8b07b43c2b3992d24314

Observation fbba7441-fa9f-4444-943d-2d1116594347 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.071076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.071076Z digest=sha256:0a650b61b0cad6962d25a4515437425999522fc5c4ba07eb0975fa8cd8e340f6

Observation 9a401f42-8a84-4a96-85aa-d93346523a38 · outbound

This paper cites Watch every step! LLM agent learning via iterative step-level process refinement.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Watch every step! LLM agent learning via iterative step-level process refinement

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.154163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.154163Z digest=sha256:0bcec8e19b224b44a7df231a87eb5674860b086726121247952ec0ec41183e9a

Observation 61f5e7dc-8d9f-4914-9e8b-0a5c6114f23a · outbound

This paper cites Watch every step! LLM agent learning via iterative step-level process refinement.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Watch every step! LLM agent learning via iterative step-level process refinement

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:08.527617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:05.239959Z digest=sha256:677c733f224d0b19cbe9e490b929ad3ac09c98f4fab19020abce0bf073b75b15

Observation 5ea0e976-666d-420d-a67b-5d90bd1ddfdc · outbound

This paper cites Qwen2.5 Technical Report.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.303985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.303985Z digest=sha256:944a450dfe46165507e91361268218ec9d87fabf79fc62d26819d096c7e3e97b

Observation 8c0eeca8-6703-4643-aacd-fdc1955a67d2 · outbound

This paper cites CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.375408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.375408Z digest=sha256:9de779d46d7b4e1619c0b3eb9b640d79b2910ed3e9998b3281084cbf549a40a8

Observation b4c4ce4d-1dc0-4c3b-9edb-93ffd23c3513 · outbound

This paper cites Narasimhan, and Yuan Cao.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Narasimhan, and Yuan Cao

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.442339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.442339Z digest=sha256:74fa6afbc258b7b208a34136d0408499b9030d524e5aece419dbefe783934e1b

Observation 31beb251-bf4e-4270-a5cf-b1ac54fdd52f · outbound

This paper cites N., Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning N., Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:08.312149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:05.501822Z digest=sha256:37926c461682beee626f44ea286836dd3a1c197515785369ca47f8af73be380d

Observation 3848eb40-f0f0-4ab3-af49-cebe1ec2661c · outbound

This paper cites Agent lumos: Unified and modular training for open-source language agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent lumos: Unified and modular training for open-source language agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.594415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.594415Z digest=sha256:1de0221073f5177f644941215d32d3b28c80d30c9812606cf47db2d0d36e4503

Observation 8084ff0b-4e90-42fa-bc50-6cd0ad92309c · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.725140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.725140Z digest=sha256:5656a37dbec9781fce60e92bf5754a7ee6c33c312f23d9e3edf4a4ee5b4e651f

Observation 60ac31f0-aca5-4678-9ad2-e093e0cffafa · outbound

This paper cites Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.817610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.817610Z digest=sha256:9a397bee0d0dad86f08aa18a226c75aa012fc42937f3b83a22a12f5e767e279b

Observation cd366988-da8b-4428-9b9e-d5d54b3591c0 · outbound

This paper cites Agenttuning: Enabling generalized agent abilities for llms.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agenttuning: Enabling generalized agent abilities for llms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.980261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.980261Z digest=sha256:7bbb0c8f7679b058fc0f52db4d0fd315601fee659358e6dd241a19b1ba4c924e

Observation ac810c3d-7994-44d2-a01e-d81a1b9d02c9 · outbound

This paper cites Enhancing decision-making for LLM agents via step-level q-value models.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Enhancing decision-making for LLM agents via step-level q-value models

Reference 58

Resolution
verified exact
doi, observed 2026-08-06T21:55:07.075035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:06.053884Z digest=sha256:82d6e24a360606891a03c3f4a9006f137e8b6f44677c5c8524f3b3deb4320a7f

Observation cfa7755b-703c-4908-98a1-9002c37c780e · outbound

This paper cites Large language models as commonsense knowledge for large-scale task planning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Large language models as commonsense knowledge for large-scale task planning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:08.111874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:06.177700Z digest=sha256:20e93bcff0cfe3b39ace89694e20894c88668e49f54f62b0e44ccda4a02e01a4

Observation c298c002-3bf0-4527-bdd5-9f85043486d1 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.305305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.305305Z digest=sha256:88b06d8d5bd939c539386def7c43a38c3d08f3a2cbca0e6dd0f7adff31dd5ebb

Observation 606cad4c-3865-416f-8d17-7281ef795196 · outbound

This paper cites Agents: An Open-source Framework for Autonomous Language Agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agents: An Open-source Framework for Autonomous Language Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.427751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.427751Z digest=sha256:50030503a5d212203520bcb2f091d1f2b07cfa2cb1e46d295fa5382e4cf78d53

Observation 5861ae12-b48c-443d-b637-12da6f2d9901 · outbound

This paper cites K now A gent: Knowledge-augmented planning for LLM -based agents.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning K now A gent: Knowledge-augmented planning for LLM -based agents

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:55:07.941110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:55:06.534526Z digest=sha256:6a1338a020a340e6b1753db40db9171e7d9b0a2cf023e2afdf4faeefd6f6f53e

Observation 61970ca3-7b6b-4135-9bf8-8e1a78af80fc · outbound

This paper cites write newline.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning write newline

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.577975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.577975Z digest=sha256:e3980af85a314a47e5545292e0550cf48010dd3744b269703c4de9b6cf7ee0fc

Observation bb9523f3-12e8-4d73-a115-2a0433ca4c7a · outbound

This paper cites @esa (Ref.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning @esa (Ref

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.675280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.675280Z digest=sha256:d00efdf1783acd8f4bde5fc30cc38a9757fe07b5c94fb7983abd33a7c91775ff

Observation 048ecf1a-75c5-496b-94d8-7d7281645b8f · outbound

This paper cites an unresolved cited work.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.732398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.732398Z digest=sha256:b3369dfc526413368aa65418e88f4d5fb5984e13714257c6c40181fc6dfbd2bf

Observation 6f087a48-75c0-4512-99f8-8e14c8b6bd0f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:06.821160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:06.821160Z digest=sha256:5f86cde26cb5cf7ab28d3e70323be4ef744fcb2ce35e1d814f23cbc5f756a71d

Pith citing papers

Observation 6f8873d8-e2c1-4d93-b663-d1c42fe201b3 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:38.485506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:38.485506Z digest=sha256:394795c8f6c02b1b411548b62d3ccd93db3830639b5adddcca7f6ac227ee49be

Observation d509ff7e-1e52-432c-8d78-a450afbf3bec · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.429654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:6618971adf093e6307de87fdc100d6d8e6f2e9f818a9c9365ebeab292a4537a5

Observation 6ce41588-75c6-48de-8f13-a86b0b9fbce5 · inbound

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning cites this paper.

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:55:33.538357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T11:52:58.630590Z digest=sha256:f5351d2b3e8f0d762fe8802df331f9d1ac2f44b77ddf563225158bf40feabcfe