Pith. sign in

Paper Citation Record · LEDGER

Reinforced Reasoning for Embodied Planning

As of 15 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 6 inbound Pith citation observations for arXiv:2505.22050.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22050 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:26.851204Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:50:32.142169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:28:15.942007Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24e9227a-ffd2-414c-8a7c-fc9cb3a023e7 · outbound

This paper cites URL: https://openai.com/index/ gpt-4o-mini-advancing-cost-efficient-intelligence/.

Reinforced Reasoning for Embodied Planning URL: https://openai.com/index/ gpt-4o-mini-advancing-cost-efficient-intelligence/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:31.068071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:20.697606Z digest=sha256:880b850e932bfc37585bf853c1a34de0e8bab4b853745aca642a8e294c286c9e

Observation 65f69c35-873e-4102-b65b-9e0a69ef96ee · outbound

This paper cites URL: https://openai.com/index/hello-gpt-4o/.

Reinforced Reasoning for Embodied Planning URL: https://openai.com/index/hello-gpt-4o/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.915929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:20.785954Z digest=sha256:27f3885be000f35c9cecaac3500c415a129d9a956bf61520f1273fbd23e72997

Observation 3b59e599-9cd7-4bf1-9ca2-cdcc395a60a3 · outbound

This paper cites URL: https://www.anthropic.com/news/ claude-3-5-sonnet.

Reinforced Reasoning for Embodied Planning URL: https://www.anthropic.com/news/ claude-3-5-sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.704191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:20.861921Z digest=sha256:909ef285792af845b4366fdfdea21b1bd27fccc251405d722bbbda6424bfaff6

Observation 3eb58784-d7ba-4e28-ad95-8d9006b35e7f · outbound

This paper cites URL: https://blog.

Reinforced Reasoning for Embodied Planning URL: https://blog

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.537405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:20.971449Z digest=sha256:19a2e582087d2b1d0b34222965e225dadc62c257a82c2de1b2465dd99ef18d7d

Observation 94ffbd43-4381-4b69-acce-c470cc9f74ba · outbound

This paper cites URL: https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/.

Reinforced Reasoning for Embodied Planning URL: https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.286481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:21.081823Z digest=sha256:a6524e67d19191748c617b927ef6f811c2c5d8a8b2ffb505eb9d732b2d838756

Observation 6d72c2fe-222c-4599-935a-6740ea66b12a · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Reinforced Reasoning for Embodied Planning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.227527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.227527Z digest=sha256:b5c828766f010a7674a6f51843e0adc9a927dd5971b01a9dcaf7eb80e5cd22b0

Observation 1466b7fc-2907-4e6d-a50a-285ceff2fbb6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Reinforced Reasoning for Embodied Planning Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.376588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.376588Z digest=sha256:945c2fbcd3b864cc155a8e4f7ead207a6ea125ed2e6970bbd814b0d688014a79

Observation 15d6ded9-411f-4a3c-8807-c8451aa847fe · outbound

This paper cites RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks.

Reinforced Reasoning for Embodied Planning RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.473781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.473781Z digest=sha256:aefb78ccea0d68cb8d3d50e98571bf261440698c984426b7ce5d8a2817332096

Observation ab1fe195-2f7d-4b5c-bb0e-a44cd8a1818e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Reinforced Reasoning for Embodied Planning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.610900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.610900Z digest=sha256:cabca754623bf9aac652862ab80ec3dfe9d7341d7e3b6af73ec9a96791705b00

Observation 7eaaf131-0b02-4d53-bea1-345e82ad6be2 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Reinforced Reasoning for Embodied Planning Process Reinforcement through Implicit Rewards

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.739893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.739893Z digest=sha256:3ac1acefb3163711d253d62b68fa9f58736a8a7cf7d2f524dc40c9b075365911

Observation ba721db9-318d-48eb-bfe2-050a9a589979 · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

Reinforced Reasoning for Embodied Planning A survey of embodied ai: From simulators to research tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.065537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:21.894396Z digest=sha256:3bda84c3c5d49ed55322d273231f0c4bfa7e901134ab975f5a41366e006b28c5

Observation 144e999e-c563-48ce-84b8-313f1e88c37d · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Reinforced Reasoning for Embodied Planning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.814108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:22.008668Z digest=sha256:d7a1a129074001da0f9c00c48b25b4b20fc7b0a8bd56b191501015df01252e0c

Observation 521e9aaa-7842-462c-aa30-ced4c358632e · outbound

This paper cites What can vlms do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition, 2024.

Reinforced Reasoning for Embodied Planning What can vlms do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.618245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:22.096292Z digest=sha256:34e5bf48214f4e38335cf6a3446ad1b80c98bbdd89e996780bac8fc403c05f10

Observation 486f722f-2e98-4710-994c-fcb43486f642 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforced Reasoning for Embodied Planning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.190695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.190695Z digest=sha256:a99f621e33ffed69b95d39936a713e5d769da02d64caa72e076afa1979fd4cb1

Observation b35a1a89-f812-4f18-8a4e-ba084039e568 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Reinforced Reasoning for Embodied Planning Lora: Low-rank adaptation of large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.314432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.314432Z digest=sha256:9f9069b6c4d02e8ae277ad0d1a36fa284aaae83f87f1c81b9db02dfdd9f7a18f

Observation f5814911-feb8-457b-8a69-998ae1b6c810 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Reinforced Reasoning for Embodied Planning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.439915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.439915Z digest=sha256:0754824240f4e9e163f61465dc3ee2d453ce79d9a354a266ddbf37abe56705f9

Observation 306cb825-6fc5-49d0-b7ac-c2cf156d7404 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

Reinforced Reasoning for Embodied Planning Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.543422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.543422Z digest=sha256:11dd7293f24ff837139d78cf4a882f65c29e8ae57034c4af29c39fbd55a7101c

Observation 953bad54-af75-45d8-a8f4-18ea783b6597 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Reinforced Reasoning for Embodied Planning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.647159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.647159Z digest=sha256:ac10fd1b8467d9247a86cc1bd69a59a5b6f3f771b6b1e637cd93fd72f74f0eef

Observation 0e8cbed3-1426-4168-89c8-6ca6e30aecaa · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Reinforced Reasoning for Embodied Planning RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.795873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.795873Z digest=sha256:43a2214b3dcb4a55e94114e512786194e55df1061cf818cf1b1ebd2bcde014a6

Observation fb0633e3-f0b5-4c7c-8c62-12fad8a46a5a · outbound

This paper cites Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents.

Reinforced Reasoning for Embodied Planning Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:22:27.092321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:22.902090Z digest=sha256:803a593cc708a6d2bfdf447bbc6272dbce932c0c1dbaaf30f57a78750e208828

Observation 015d696c-6ac0-4202-80c8-9d98520e52c0 · outbound

This paper cites Openvla: An open-source vision-language-action model.

Reinforced Reasoning for Embodied Planning Openvla: An open-source vision-language-action model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.432336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:23.000783Z digest=sha256:3ee9ed3d01119b324284c3e6b27896da22ce25f446d48f7a4168a6615d25451c

Observation 59e6988c-2cf4-4ed4-b91d-b3763ddf060e · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Reinforced Reasoning for Embodied Planning AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.067075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.067075Z digest=sha256:69051c47379384f266d27db85f89ebce112b40fb8ece5cf028f54d7b9265f30d

Observation 0a9a00ae-3fd0-46c2-be4e-a04520f0240e · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Reinforced Reasoning for Embodied Planning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.131565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.131565Z digest=sha256:eabc0b5dcefa4704391af9e681b9955dcc0a09dbc44cf52323d084dab7bbf0fa

Observation cd364a8c-639e-4932-8fb5-0d90a571eece · outbound

This paper cites Let’s verify step by step.

Reinforced Reasoning for Embodied Planning Let’s verify step by step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.205810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.205810Z digest=sha256:66afa7236e1a15641f0e96562493e28024cb10d96ae30f0dd04f6d7f88cf4762

Observation 068979f8-27f0-45ce-abe9-af07d4cc13f9 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Reinforced Reasoning for Embodied Planning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.320295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.320295Z digest=sha256:26a4764958cdd0b07e406229bfe84f05d86f5fe0578184c4ccd24dcc32ff5aaa

Observation 75ed4314-1218-4b26-9220-0e431c3a42f3 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

Reinforced Reasoning for Embodied Planning A Survey on Vision-Language-Action Models for Embodied AI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.425995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.425995Z digest=sha256:7e5ea53fff0cbfa2618e6e47325846baa1dbba3f3331c605d266567efb12ae6e

Observation 594ba930-4cde-4303-9a9a-16a9b0917d39 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Reinforced Reasoning for Embodied Planning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.503690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.503690Z digest=sha256:daf1118dae03e704f1580c89eebfe9a49cac85bae607d5ff88840596cc3d1fb3

Observation c4064ed7-2bcd-4381-970d-e4a9ac66cc24 · outbound

This paper cites Composi- tional chain-of-thought prompting for large multimodal models.

Reinforced Reasoning for Embodied Planning Composi- tional chain-of-thought prompting for large multimodal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.284038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:23.564040Z digest=sha256:6738013c544f80f1f548dbf576e5a6d2502b810fa10621042e8238a621d0ef70

Observation e6233108-f2b8-4000-acc6-16e29d635747 · outbound

This paper cites Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning.

Reinforced Reasoning for Embodied Planning Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.088642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:23.666632Z digest=sha256:2cfbc554bbc5b8ffc5230df1d5a76c01cf6383fe182e35cfc8ab9c25762318e3

Observation a379916d-3750-496e-bd5e-d1b9bb0b6812 · outbound

This paper cites EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought.

Reinforced Reasoning for Embodied Planning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.750410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.750410Z digest=sha256:a822f451463e00152ed5444bbc74e36afec8bd88f01c68e2cc4575aac070da02

Observation 2893fc7d-e8dd-4666-bbd9-bb90b0675797 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforced Reasoning for Embodied Planning Training language models to follow instructions with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.820276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.820276Z digest=sha256:a779156315db5c132c2cc9773d4a0d4452c2e88676fffbb60e549ba9e4aa41d1

Observation b6b71b1c-ea92-4cd9-9494-f9eafb0d7505 · outbound

This paper cites Reasoning with large language models, a survey.

Reinforced Reasoning for Embodied Planning Reasoning with large language models, a survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.927663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.927663Z digest=sha256:ff11d84205f76a87c52d88b98112e631eebfd50af7c1be23ae2814405bc57322

Observation d138363b-8b9b-4f20-9d31-e6b264d3761b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Reinforced Reasoning for Embodied Planning Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.070429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.070429Z digest=sha256:93d647b57bd43c8de6d6f64f76766f28a7e3e11658f0e498515d54b50c050864

Observation 2ae79c2a-828f-4a90-8707-f0c4835ee5d8 · outbound

This paper cites Say- Plan: Grounding large language models using 3d scene graphs for scalable robot task planning.

Reinforced Reasoning for Embodied Planning Say- Plan: Grounding large language models using 3d scene graphs for scalable robot task planning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.955807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:24.159911Z digest=sha256:b5a60e52f444e04b5d3f7ce2d07172e2231eab805e28fd494c14e52950abb208

Observation 2673d062-04d3-4f5a-b57a-7fce5fcb0cac · outbound

This paper cites Habitat: A platform for embodied ai research.

Reinforced Reasoning for Embodied Planning Habitat: A platform for embodied ai research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.234147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.234147Z digest=sha256:6cf83ed3415d60cddfc2e9d74bf8ed5551f940d1d028e2c79b966ebe6d5eff93

Observation d065253d-040a-4b08-a27e-ec1e2567bcfe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforced Reasoning for Embodied Planning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.317114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.317114Z digest=sha256:6a97e1743acdf621394f861a4f4637c4eb646a76652771ab76bac26a1555f9d4

Observation 82715e8d-d392-48f3-ab1c-027a5694d301 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Reinforced Reasoning for Embodied Planning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.441874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.441874Z digest=sha256:7fc466e8f83b11665267597f1dd0b0f8f70d8c58c4310799173089b5657299f7

Observation 695f2968-25c7-4f31-8cea-ac6239232391 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Reinforced Reasoning for Embodied Planning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.515252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.515252Z digest=sha256:00ff0bc2227d7fe9b08d0a1a12f5810c35107c96c0b745fec2ea317e6830c2d9

Observation ccb339cc-995d-455f-be9d-1b86e9eb32e4 · outbound

This paper cites Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following.

Reinforced Reasoning for Embodied Planning Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:22:27.751720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:24.564479Z digest=sha256:c8158fa549d0060f7a25d5db96940a5520f5e5b0470a4f59a548ac4878e1ac5a

Observation e46ce3b6-0470-4a02-a8e7-14900943e03c · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Reinforced Reasoning for Embodied Planning Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.729821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:24.632852Z digest=sha256:5a21f166e59600bff11b7d084f3f04bf9c926f7d874cd861bd6f058429116dfb

Observation 9082b7a0-ca47-426e-8226-73f21312d533 · outbound

This paper cites Tenenbaum, Leslie Kaelbling, and Michael Katz.

Reinforced Reasoning for Embodied Planning Tenenbaum, Leslie Kaelbling, and Michael Katz

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.721896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.721896Z digest=sha256:62451ff24b47ba76dd8c30e8b0f9bd3544c6b44d147cf5281f7feaaabc8990b4

Observation 6c45aed3-41bc-4ada-bbe9-090799f28ba2 · outbound

This paper cites ProgPrompt: Generating Situated Robot Task Plans using Large Language Models.

Reinforced Reasoning for Embodied Planning ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.847360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.847360Z digest=sha256:a4b0ca1100da40cb54fadab8d376fa9b278033e9b1f7b2a44ba643d75d3b5e40

Observation 5083a2e4-4759-4062-869c-bc31d46e0826 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

Reinforced Reasoning for Embodied Planning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.912504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.912504Z digest=sha256:d6e91514f16775d9dcd7b1298309db319a45765c8935cf96e5fd704300b003c9

Observation 276bf8bd-f071-4bbf-b57f-362978007a35 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Reinforced Reasoning for Embodied Planning Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.024255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.024255Z digest=sha256:75b91adb765973480768a2149a905a30966d84e9a3f9333992d6c4b2e87b2df5

Observation 5a330e41-731c-4acf-98b7-a7e1e58d8eb7 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reinforced Reasoning for Embodied Planning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.118215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.118215Z digest=sha256:5d40e9bc2dc1aafbbe801613ca344631ffad1ad7a9e505af20c97a47bf078dbf

Observation 1ef661ec-05e2-4e3e-be61-65c46d171ec2 · outbound

This paper cites World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning.

Reinforced Reasoning for Embodied Planning World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.221513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.221513Z digest=sha256:fcfee2f3a6a2ad21458ae6c6d946479563adb143c0dc3a5aa604f87e0fd76347

Observation ea7a803f-cf6c-450e-8f42-ca62b9577c92 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Reinforced Reasoning for Embodied Planning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.297429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.297429Z digest=sha256:5dad5dd6c28629e568129462e9a8c0ad7443a3eeaea881198a25b7fe90056f1c

Observation 41ebaff6-6b3f-4a3f-b640-c6982e3f7ce5 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

Reinforced Reasoning for Embodied Planning Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.358034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.358034Z digest=sha256:af270e151b304f51618e9b647a04ada3207e117d8634905713fabff56787f89c

Observation fc5ff3e5-e4fe-45b7-9a14-d9e94a21e443 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Reinforced Reasoning for Embodied Planning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.429370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.429370Z digest=sha256:b5ce03f6c8e74aff5eb7ad29e3c36d72ef7f89ebc29c7454cd159c73b22b152c

Observation 15fc3c6d-8b35-4c25-836b-2832a0b31c7f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Reinforced Reasoning for Embodied Planning Chain-of-thought prompting elicits reasoning in large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.565706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.565706Z digest=sha256:a8149f970350decd025ff7ba5759889d59fb736b18d34eed68f70107f0d1076f

Observation 5a3083d7-dbd6-4f84-9679-685ec544d6eb · outbound

This paper cites Embodied Task Planning with Large Language Models.

Reinforced Reasoning for Embodied Planning Embodied Task Planning with Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.663807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.663807Z digest=sha256:10a1aa0e42a90a21cdd77982672ed844f679a8f60264842dd503e025f29614a6

Observation 4b9d0635-0905-4b8f-a6d7-e53cb7270b75 · outbound

This paper cites The rise and potential of large language model based agents: A survey.

Reinforced Reasoning for Embodied Planning The rise and potential of large language model based agents: A survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.727513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.727513Z digest=sha256:15158230f12d6703ef48c8436f2267a0c565dc442337d5ce37163b885adb244f

Observation cf71a2cf-9922-41d1-8c85-ac16f9273565 · outbound

This paper cites A Survey on Robotics with Foundation Models: toward Embodied AI.

Reinforced Reasoning for Embodied Planning A Survey on Robotics with Foundation Models: toward Embodied AI

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.853615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.853615Z digest=sha256:8ae9f7571cbe6884b5f0b0556ff17cd52f19c4341c566d94b0fbcfb180c164dc

Observation 33ed233e-2815-4edb-9741-68178e0e3971 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Reinforced Reasoning for Embodied Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.974006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.974006Z digest=sha256:039e42fc66dc500f5fd1acc9d327bc51680db7fd66d8d318c23b9513cac64d3b

Observation b0ed911a-61c7-49ab-87f3-87547d71b64f · outbound

This paper cites LIMO: Less is More for Reasoning.

Reinforced Reasoning for Embodied Planning LIMO: Less is More for Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.046698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.046698Z digest=sha256:bf615baa3e7a2e3c00591cbaa7492610d45135ebd376ef27ae3812bc4a1ac777

Observation d3d90c58-7546-4690-99e6-d95fd96565d2 · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

Reinforced Reasoning for Embodied Planning Robotic control via embodied chain-of-thought reasoning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.569347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:26.135876Z digest=sha256:4637e562b955b97859133185a2c2c49097f8933fa01b73fd5962100f402f5743

Observation 7b573dbf-8950-48cd-911b-78c2857b79dd · outbound

This paper cites HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers.

Reinforced Reasoning for Embodied Planning HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.186961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.186961Z digest=sha256:efde0ce752047458802fba0326d84d3cc11e54f3bad013fd928b171ec8c733d7

Observation e896157e-17c9-47f5-bca2-da733487fbde · outbound

This paper cites Vision-language models for vision tasks: A survey.

Reinforced Reasoning for Embodied Planning Vision-language models for vision tasks: A survey

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.234405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.234405Z digest=sha256:61b7e50611775d224d5190faff05029a5e5cbe121ed5d3e68252a50ea68626bc

Observation 2401c5a6-982c-49ee-81e2-e35b5de658ad · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Reinforced Reasoning for Embodied Planning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.330503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.330503Z digest=sha256:c2a1cf9c268744d5da431b64e394e50cd9622da50b4215ad971e44d6d5399e6e

Observation 3365ba74-cedf-4a8a-bb6a-1d6134f17680 · outbound

This paper cites Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks.

Reinforced Reasoning for Embodied Planning Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.448609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.448609Z digest=sha256:e39556c1a98238113c00aabff6712ad05ecc46f2cf22370b2299a0603a97c2e9

Observation ba02ac77-4754-4ff3-8aeb-e459fef570f7 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Reinforced Reasoning for Embodied Planning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.520134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.520134Z digest=sha256:af1336c7332512e5ef378bda68eb7b879d179d2e69565b1458020b7cf31a7aef

Observation 72e92544-ffec-44c3-805d-e83ce60e3698 · outbound

This paper cites Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning.

Reinforced Reasoning for Embodied Planning Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.571212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.571212Z digest=sha256:9d21c4072790588c766be69f500ab9201ec36b5464bbc37b2cb1c5acd372c71a

Observation d2b29aa4-dbbc-4fdb-8da3-7385faca2cff · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Reinforced Reasoning for Embodied Planning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:22:26.635776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.635776Z digest=sha256:09876d37064b11d4f020921e4de882fbdcc0c9443a4d98cc4ecfd14c420feaee

Observation 1f57eff8-5fb6-4f00-974e-911afd275ee8 · outbound

This paper cites reasoning_and_reflection\.

Reinforced Reasoning for Embodied Planning reasoning_and_reflection\

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.353645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:26.700111Z digest=sha256:1073f8dcae713527ee364965ef1ca3ac9780347b4d1602cea648fdb24422721a

Observation 9bede566-deae-4e2d-a0ba-ca8adacd5eec · outbound

This paper cites an unresolved cited work.

Reinforced Reasoning for Embodied Planning Unresolved cited work

Reference 224

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:28.207679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:26.851204Z digest=sha256:a6a47da72825883f9c80b417ceaaa4468e7797bcef483f6a3cb798e0e0c133f6

Pith citing papers

Observation dc9d4a05-26d6-4f64-b7e6-b01053167dd8 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Reinforced Reasoning for Embodied Planning

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:15.947442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:13b9cc2419d674ea9211ac1a593cf5e941d07f00ef7198f338c2fc1baae3f403

Observation 8355abd1-3ad3-4a41-b036-c0dc8ba2d1d3 · inbound

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial cites this paper.

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial Reinforced Reasoning for Embodied Planning

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:32.142169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:32.142169Z digest=sha256:bc7d6d53151f26f635df4de59e096c2e6adf0b91f3efd290587413d5647d6ff6

Observation 228d765b-1279-480a-b96e-82f25ef54e93 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Reinforced Reasoning for Embodied Planning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:41.474386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:41.474386Z digest=sha256:cbc04d42d0235cb9b7484601b45cc26024a69ed07ba08c349345357b4749e677

Observation d0aefcf2-b090-463b-a156-431113c66b6c · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Reinforced Reasoning for Embodied Planning

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.247031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:4f53469ed45d8791dc210beb281c5040aa68721f0c66db7fd0c04d4e0d01ec6e

Observation 29bc1a9f-64c8-42e9-b5c6-1f94a5eab6b9 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning Reinforced Reasoning for Embodied Planning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:09.217964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:aca8967895143aeaa028681d796ff0a006e344d238e042687ea1d8d6f6bffe16

Observation d7ee1b53-19af-41f1-91ac-da99f4e9d0b8 · inbound

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data cites this paper.

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data Reinforced Reasoning for Embodied Planning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:57:33.219109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T17:54:50.325820Z digest=sha256:2a30a07c0570aa1906beb4d734276e2d627186e1139ad0f90a38f94ae024225d