Pith. sign in

Paper Citation Record · LEDGER

Reinforced Reasoning for Embodied Planning

As of 10 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 6 inbound Pith citation observations for arXiv:2505.22050.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22050 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:26.851204Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:50:32.142169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:28:15.942007Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24e9227a-ffd2-414c-8a7c-fc9cb3a023e7 · outbound

This paper cites URL: https://openai.com/index/ gpt-4o-mini-advancing-cost-efficient-intelligence/.

Reinforced Reasoning for Embodied Planning URL: https://openai.com/index/ gpt-4o-mini-advancing-cost-efficient-intelligence/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:31.068071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:20.697606Z digest=sha256:0433209f5459981cf007c35dd062226c4bcdbfe675f7274101f3e2166e4f4f17

Observation 65f69c35-873e-4102-b65b-9e0a69ef96ee · outbound

This paper cites URL: https://openai.com/index/hello-gpt-4o/.

Reinforced Reasoning for Embodied Planning URL: https://openai.com/index/hello-gpt-4o/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.915929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:20.785954Z digest=sha256:83dead19ca750cd3bcec5bf1066f39789684410473a9f1597b98e0f850bd138e

Observation 3b59e599-9cd7-4bf1-9ca2-cdcc395a60a3 · outbound

This paper cites URL: https://www.anthropic.com/news/ claude-3-5-sonnet.

Reinforced Reasoning for Embodied Planning URL: https://www.anthropic.com/news/ claude-3-5-sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.704191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:20.861921Z digest=sha256:5112b6b166b43e2cc3c3890f9f6201fa312bcce652fae9b54260c1af8db8ea01

Observation 3eb58784-d7ba-4e28-ad95-8d9006b35e7f · outbound

This paper cites URL: https://blog.

Reinforced Reasoning for Embodied Planning URL: https://blog

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.537405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:20.971449Z digest=sha256:05a438dcbd5b003ad4b597fdc37c46de6428002afe94708836d1cf8f7129ed0f

Observation 94ffbd43-4381-4b69-acce-c470cc9f74ba · outbound

This paper cites URL: https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/.

Reinforced Reasoning for Embodied Planning URL: https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.286481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:21.081823Z digest=sha256:bfef539c416c1f55ae5d159b6e19ad69b40d6151f16c7d2eac01ae9e690b0e24

Observation 6d72c2fe-222c-4599-935a-6740ea66b12a · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Reinforced Reasoning for Embodied Planning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.227527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.227527Z digest=sha256:ff5d3551899cbf82b3eeb1c57a8e9741c602c85b1c45d179fface3e92d9f67c4

Observation 1466b7fc-2907-4e6d-a50a-285ceff2fbb6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Reinforced Reasoning for Embodied Planning Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.376588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.376588Z digest=sha256:c6e04f2a108fa400e42fceeac1495bb94fcb19a3ab11158ded2835ecdd0d7703

Observation 15d6ded9-411f-4a3c-8807-c8451aa847fe · outbound

This paper cites RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks.

Reinforced Reasoning for Embodied Planning RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.473781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.473781Z digest=sha256:413eabc089a437d822636e1a6cab51c9e6b0703a352554ee3f06a1a0eeec3bc2

Observation ab1fe195-2f7d-4b5c-bb0e-a44cd8a1818e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Reinforced Reasoning for Embodied Planning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.610900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.610900Z digest=sha256:4873f80de52bbb3eb3ce16de6ababe48b727054e0287fefb449b7517305566ce

Observation 7eaaf131-0b02-4d53-bea1-345e82ad6be2 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Reinforced Reasoning for Embodied Planning Process Reinforcement through Implicit Rewards

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.739893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.739893Z digest=sha256:a599ded32182554079439b95db8be3f828258abd419e316bbe7601df0d747590

Observation ba721db9-318d-48eb-bfe2-050a9a589979 · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

Reinforced Reasoning for Embodied Planning A survey of embodied ai: From simulators to research tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.065537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:21.894396Z digest=sha256:a87070d39814a4a5f3464d6c32bc148f841aa0b5b2d92e948e13eba5a89c8c4c

Observation 144e999e-c563-48ce-84b8-313f1e88c37d · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Reinforced Reasoning for Embodied Planning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.814108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:22.008668Z digest=sha256:0b8ff01973fe1f08ba44a8c39023aa62cb00e47f7a7d3114a26773bc3904aa16

Observation 521e9aaa-7842-462c-aa30-ced4c358632e · outbound

This paper cites What can vlms do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition, 2024.

Reinforced Reasoning for Embodied Planning What can vlms do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.618245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:22.096292Z digest=sha256:d9758da1cbd6772120b0e058783a3b265758d7a8f4c495f61468a11a01a2cf33

Observation 486f722f-2e98-4710-994c-fcb43486f642 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforced Reasoning for Embodied Planning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.190695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.190695Z digest=sha256:3499dcd40d5f4a5bf575d22ae0833ac169f12395f2411900041d04a3aacab9e1

Observation b35a1a89-f812-4f18-8a4e-ba084039e568 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Reinforced Reasoning for Embodied Planning Lora: Low-rank adaptation of large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.314432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.314432Z digest=sha256:3826ef806a613232550952eac0bdbfb4f2e523015a68f9d27df4be82625b8057

Observation f5814911-feb8-457b-8a69-998ae1b6c810 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Reinforced Reasoning for Embodied Planning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.439915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.439915Z digest=sha256:e3ae6ef57bf90dc7fa7b47cf0ac3f69a9041de75792ff3aefb9414b555e3242c

Observation 306cb825-6fc5-49d0-b7ac-c2cf156d7404 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

Reinforced Reasoning for Embodied Planning Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.543422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.543422Z digest=sha256:0c26a510b448edcddb54b4d0059067531598d7f619805f623e16ac5b72c2818c

Observation 953bad54-af75-45d8-a8f4-18ea783b6597 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Reinforced Reasoning for Embodied Planning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.647159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.647159Z digest=sha256:8fa0428f590051809508ed0a092575fcb11558fd23a318832dbf31144f7fc05c

Observation 0e8cbed3-1426-4168-89c8-6ca6e30aecaa · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Reinforced Reasoning for Embodied Planning RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.795873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.795873Z digest=sha256:bd1dd3d0c19acd07c77697420fd4c5cb0ddb539a9e3972fb690dcdde3eeb810c

Observation fb0633e3-f0b5-4c7c-8c62-12fad8a46a5a · outbound

This paper cites Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents.

Reinforced Reasoning for Embodied Planning Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:22:27.092321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:22.902090Z digest=sha256:01360d181ff17c3af9ae1ffd46e684f80cbe479464c92d10bfe118bb11d95304

Observation 015d696c-6ac0-4202-80c8-9d98520e52c0 · outbound

This paper cites Openvla: An open-source vision-language-action model.

Reinforced Reasoning for Embodied Planning Openvla: An open-source vision-language-action model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.432336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:23.000783Z digest=sha256:e96f976378e0da150a65e4d7b27219ed9ad9aa159b93aead3852cb2f1314f6da

Observation 59e6988c-2cf4-4ed4-b91d-b3763ddf060e · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Reinforced Reasoning for Embodied Planning AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.067075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.067075Z digest=sha256:4cfdbb354cc6de93b822c11ab10722437d68f8577bff88634748eb657cd2d783

Observation 0a9a00ae-3fd0-46c2-be4e-a04520f0240e · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Reinforced Reasoning for Embodied Planning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.131565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.131565Z digest=sha256:7472dcf8b7270df96a7652bff1294e2a5b188431ba2c10af07f4c3e32d89aa0d

Observation cd364a8c-639e-4932-8fb5-0d90a571eece · outbound

This paper cites Let’s verify step by step.

Reinforced Reasoning for Embodied Planning Let’s verify step by step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.205810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.205810Z digest=sha256:020c9ac738de68b9f4a3eefb65bb1f9209109d67d53f33745a2fb369e0adf5d2

Observation 068979f8-27f0-45ce-abe9-af07d4cc13f9 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Reinforced Reasoning for Embodied Planning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.320295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.320295Z digest=sha256:81a54f707d43b46a27820976ed928c7e695ea8f30470d5b10db4e673f340615a

Observation 75ed4314-1218-4b26-9220-0e431c3a42f3 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

Reinforced Reasoning for Embodied Planning A Survey on Vision-Language-Action Models for Embodied AI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.425995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.425995Z digest=sha256:1a01ce8609f04f2acdef4a6f5c7f54a0b8ef1e4f52b2ffb858bebf6bf11b68d8

Observation 594ba930-4cde-4303-9a9a-16a9b0917d39 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Reinforced Reasoning for Embodied Planning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.503690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.503690Z digest=sha256:2e7cf58b854d58b8e5440499e8e721b48eb06d2f797d62c0ffad9fe4ab4ef04e

Observation c4064ed7-2bcd-4381-970d-e4a9ac66cc24 · outbound

This paper cites Composi- tional chain-of-thought prompting for large multimodal models.

Reinforced Reasoning for Embodied Planning Composi- tional chain-of-thought prompting for large multimodal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.284038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:23.564040Z digest=sha256:3adcaf29efc1da69deebb58d46d3c3d748e06496d7784b57808b9c3ec490456b

Observation e6233108-f2b8-4000-acc6-16e29d635747 · outbound

This paper cites Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning.

Reinforced Reasoning for Embodied Planning Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.088642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:23.666632Z digest=sha256:ce2dcb3853d187254c05e0f7aa1a8adccb6923cc82ad4f8c6f2a1eb93579ed64

Observation a379916d-3750-496e-bd5e-d1b9bb0b6812 · outbound

This paper cites EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought.

Reinforced Reasoning for Embodied Planning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.750410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.750410Z digest=sha256:2b0badcdac98b3182233933dce450cc6ca6a93d9daaeb8c4112a5fab60900119

Observation 2893fc7d-e8dd-4666-bbd9-bb90b0675797 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforced Reasoning for Embodied Planning Training language models to follow instructions with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.820276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.820276Z digest=sha256:4c1bec33c06e5c116e78dd42f728af9bfd47a560c267cada842b78ad824cad16

Observation b6b71b1c-ea92-4cd9-9494-f9eafb0d7505 · outbound

This paper cites Reasoning with large language models, a survey.

Reinforced Reasoning for Embodied Planning Reasoning with large language models, a survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.927663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.927663Z digest=sha256:ea714e7186119820a19255b69a8d6c17900788e20df75b1ea35bac7406110b20

Observation d138363b-8b9b-4f20-9d31-e6b264d3761b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Reinforced Reasoning for Embodied Planning Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.070429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.070429Z digest=sha256:177cb673059e5ab8df1ec55fa080fcb6d25ff25fc977995301721c63ca02e474

Observation 2ae79c2a-828f-4a90-8707-f0c4835ee5d8 · outbound

This paper cites Say- Plan: Grounding large language models using 3d scene graphs for scalable robot task planning.

Reinforced Reasoning for Embodied Planning Say- Plan: Grounding large language models using 3d scene graphs for scalable robot task planning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.955807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:24.159911Z digest=sha256:6502e1af63a7a3b640dd81e3c8f154361160e3e9000e67bdd9461d08cd23b3f1

Observation 2673d062-04d3-4f5a-b57a-7fce5fcb0cac · outbound

This paper cites Habitat: A platform for embodied ai research.

Reinforced Reasoning for Embodied Planning Habitat: A platform for embodied ai research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.234147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.234147Z digest=sha256:48aa44e3c6f781f87f3fca4d967889ef931329ae1ad868d5e4b651fe09cf3bdf

Observation d065253d-040a-4b08-a27e-ec1e2567bcfe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforced Reasoning for Embodied Planning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.317114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.317114Z digest=sha256:aadbf594a0b2c4c73d40cbea428f5087fe5eb9b4cadca52dd300d007f74e9d5d

Observation 82715e8d-d392-48f3-ab1c-027a5694d301 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Reinforced Reasoning for Embodied Planning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.441874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.441874Z digest=sha256:13fb4132d9c5c23430b0342b89ce3be11d05b4106d0413c88eefe44acabaaf40

Observation 695f2968-25c7-4f31-8cea-ac6239232391 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Reinforced Reasoning for Embodied Planning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.515252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.515252Z digest=sha256:d4181dde283421e74b57eca6b9746e2eb66eca9d02a2bd4b8c7e29c1648a17e8

Observation ccb339cc-995d-455f-be9d-1b86e9eb32e4 · outbound

This paper cites Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following.

Reinforced Reasoning for Embodied Planning Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:22:27.751720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:24.564479Z digest=sha256:ca00048538dd1569bee83b515536644967bca01335f06acc91ae404c706459e3

Observation e46ce3b6-0470-4a02-a8e7-14900943e03c · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Reinforced Reasoning for Embodied Planning Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.729821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:24.632852Z digest=sha256:4db2f035edf386d6e274d9d123928d4f54f73255e5a27054ee46e9826d36d27d

Observation 9082b7a0-ca47-426e-8226-73f21312d533 · outbound

This paper cites Tenenbaum, Leslie Kaelbling, and Michael Katz.

Reinforced Reasoning for Embodied Planning Tenenbaum, Leslie Kaelbling, and Michael Katz

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.721896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.721896Z digest=sha256:c31ec9a9ed9ef663c576b880149579a721545401158bd6d241be37f2794d19a1

Observation 6c45aed3-41bc-4ada-bbe9-090799f28ba2 · outbound

This paper cites ProgPrompt: Generating Situated Robot Task Plans using Large Language Models.

Reinforced Reasoning for Embodied Planning ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.847360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.847360Z digest=sha256:9cdec1fa33530cc500517ebd245250d87df5cecf986b8bfb03e4352654678b0a

Observation 5083a2e4-4759-4062-869c-bc31d46e0826 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

Reinforced Reasoning for Embodied Planning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.912504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.912504Z digest=sha256:2379b1e98f4d7ea5c9b7f1439d974ef451c1594e7f9fcfd83ece57628fe85c45

Observation 276bf8bd-f071-4bbf-b57f-362978007a35 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Reinforced Reasoning for Embodied Planning Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.024255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.024255Z digest=sha256:902a20e19c455e42906d25c581bda8dbdb871cc4d3f160b328f405705a750cb9

Observation 5a330e41-731c-4acf-98b7-a7e1e58d8eb7 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reinforced Reasoning for Embodied Planning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.118215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.118215Z digest=sha256:6ef5fef573f023dabbbab36645bfd1f5298569cbc22b5c8e9bd38b2a3b3de418

Observation 1ef661ec-05e2-4e3e-be61-65c46d171ec2 · outbound

This paper cites World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning.

Reinforced Reasoning for Embodied Planning World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.221513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.221513Z digest=sha256:5e59a78ed13a4ad4258565029957d42b675e71a0de42d3b3f2dc44a343164983

Observation ea7a803f-cf6c-450e-8f42-ca62b9577c92 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Reinforced Reasoning for Embodied Planning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.297429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.297429Z digest=sha256:5c1c3d6208785c4b53844e58da8fcff8011616fb319339b84184ff8331b255f3

Observation 41ebaff6-6b3f-4a3f-b640-c6982e3f7ce5 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

Reinforced Reasoning for Embodied Planning Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.358034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.358034Z digest=sha256:78964720e0465fe8d1afde6d72de14ed249bdc04c023308e535385d6cc193a5d

Observation fc5ff3e5-e4fe-45b7-9a14-d9e94a21e443 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Reinforced Reasoning for Embodied Planning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.429370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.429370Z digest=sha256:37d942855efe90c1ffeef239b02b95c74650591448e3b24f35f6b4ef9956180b

Observation 15fc3c6d-8b35-4c25-836b-2832a0b31c7f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Reinforced Reasoning for Embodied Planning Chain-of-thought prompting elicits reasoning in large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.565706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.565706Z digest=sha256:64d8233bade711e33f06a7fbdad0a0da5f29b4b3886534f3fd5b7319a1e32d14

Observation 5a3083d7-dbd6-4f84-9679-685ec544d6eb · outbound

This paper cites Embodied Task Planning with Large Language Models.

Reinforced Reasoning for Embodied Planning Embodied Task Planning with Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.663807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.663807Z digest=sha256:45c8efc469d2190373e111156b815ddf3b0f501aa1d461c5dc751c7515e3405d

Observation 4b9d0635-0905-4b8f-a6d7-e53cb7270b75 · outbound

This paper cites The rise and potential of large language model based agents: A survey.

Reinforced Reasoning for Embodied Planning The rise and potential of large language model based agents: A survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.727513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.727513Z digest=sha256:d8a457902acf9803dfead28c566f1eedaa26266725a44cace4038833809eeeb3

Observation cf71a2cf-9922-41d1-8c85-ac16f9273565 · outbound

This paper cites A Survey on Robotics with Foundation Models: toward Embodied AI.

Reinforced Reasoning for Embodied Planning A Survey on Robotics with Foundation Models: toward Embodied AI

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.853615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.853615Z digest=sha256:39a11820b5f588dfdf005fb978a4d5cb2d2ed2995c0b56bf0ddfeda14e2dad92

Observation 33ed233e-2815-4edb-9741-68178e0e3971 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Reinforced Reasoning for Embodied Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.974006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.974006Z digest=sha256:0d135e4698f827d9a8c3568ced07b213a468066105e8d9e44416cd918720b285

Observation b0ed911a-61c7-49ab-87f3-87547d71b64f · outbound

This paper cites LIMO: Less is More for Reasoning.

Reinforced Reasoning for Embodied Planning LIMO: Less is More for Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.046698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.046698Z digest=sha256:b768c6d633ba41f1cc859bc1731b3a9a0f305e9d04fd95f5fd33521bc534c350

Observation d3d90c58-7546-4690-99e6-d95fd96565d2 · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

Reinforced Reasoning for Embodied Planning Robotic control via embodied chain-of-thought reasoning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.569347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:26.135876Z digest=sha256:c4c088d862774950c30aa49edd95f54d439b5517895ca6f97f6f0cfe28d49564

Observation 7b573dbf-8950-48cd-911b-78c2857b79dd · outbound

This paper cites HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers.

Reinforced Reasoning for Embodied Planning HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.186961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.186961Z digest=sha256:254e10b99e39480617388b6199c717a2ef1649b853e5858d502b68541f1cbdb4

Observation e896157e-17c9-47f5-bca2-da733487fbde · outbound

This paper cites Vision-language models for vision tasks: A survey.

Reinforced Reasoning for Embodied Planning Vision-language models for vision tasks: A survey

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.234405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.234405Z digest=sha256:4fd19a9522e4979ad5bdc77f32839fc1a2615470b4b2a827be3e78521e5955a8

Observation 2401c5a6-982c-49ee-81e2-e35b5de658ad · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Reinforced Reasoning for Embodied Planning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.330503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.330503Z digest=sha256:bd4cfab64713b5337bb8a5371729390dc63ba8ebb126c77cbbd0716638b4b3a1

Observation 3365ba74-cedf-4a8a-bb6a-1d6134f17680 · outbound

This paper cites Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks.

Reinforced Reasoning for Embodied Planning Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.448609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.448609Z digest=sha256:3aff4bf8af95b6b95a9932e305b23ea1aa3c67c243ebeba44368c2271b74f6c4

Observation ba02ac77-4754-4ff3-8aeb-e459fef570f7 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Reinforced Reasoning for Embodied Planning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.520134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.520134Z digest=sha256:8e5c33df35983e43aa4105c9aa6de756363cbbcd4dfc378ec3be4c6598791d30

Observation 72e92544-ffec-44c3-805d-e83ce60e3698 · outbound

This paper cites Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning.

Reinforced Reasoning for Embodied Planning Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.571212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.571212Z digest=sha256:3c20c8d83a70e879af6f9121f65d24620d9b8ba843d17b72e0409dc0f49d3099

Observation d2b29aa4-dbbc-4fdb-8da3-7385faca2cff · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Reinforced Reasoning for Embodied Planning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:22:26.635776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.635776Z digest=sha256:eec4d0549d922804f389e4bee30eb6f7eaa497b791759c37a0d07d0d31a9c3de

Observation 1f57eff8-5fb6-4f00-974e-911afd275ee8 · outbound

This paper cites reasoning_and_reflection\.

Reinforced Reasoning for Embodied Planning reasoning_and_reflection\

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.353645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:26.700111Z digest=sha256:3d223555dda0eddee69b824e930e2c027d89d8c1a2fe5961793c50acdcb2cf60

Observation 9bede566-deae-4e2d-a0ba-ca8adacd5eec · outbound

This paper cites an unresolved cited work.

Reinforced Reasoning for Embodied Planning Unresolved cited work

Reference 224

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:28.207679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:22:26.851204Z digest=sha256:5dcc6604b4821180d31f021e01973b0c343a56d82ec90cbf38ef0e2879cc231a

Pith citing papers

Observation dc9d4a05-26d6-4f64-b7e6-b01053167dd8 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Reinforced Reasoning for Embodied Planning

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:15.947442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:c14fdd251edb13d745a84efcd0381ce1be36fdcdbc11b98987402c58cbab4e6b

Observation 8355abd1-3ad3-4a41-b036-c0dc8ba2d1d3 · inbound

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial cites this paper.

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial Reinforced Reasoning for Embodied Planning

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:32.142169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:32.142169Z digest=sha256:76fb0358ea92a48911080b37910a9b54025cf5a6491f3564ae607f0010445a35

Observation 228d765b-1279-480a-b96e-82f25ef54e93 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Reinforced Reasoning for Embodied Planning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:41.474386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:41.474386Z digest=sha256:79a362a0cda640956835d644f1ff049d308bce646d63f70274646603a9aac450

Observation d0aefcf2-b090-463b-a156-431113c66b6c · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Reinforced Reasoning for Embodied Planning

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.247031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:0a1ba16495d2b247011df70fbfbecedca99f2ca8946f7f7d03a39efcc79f9cb1

Observation 29bc1a9f-64c8-42e9-b5c6-1f94a5eab6b9 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning Reinforced Reasoning for Embodied Planning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:09.217964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:bf528931126f7ca2eda60b68af82598c14203e4b156c94cb3c7c8ddabe0bae5f

Observation d7ee1b53-19af-41f1-91ac-da99f4e9d0b8 · inbound

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data cites this paper.

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data Reinforced Reasoning for Embodied Planning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:57:33.219109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T17:54:50.325820Z digest=sha256:1b025cdc3ad91090974d656e98d497161cf1be55cea62f3adce56c7aa28f9f6c