Pith. sign in

REVIEW 9 cited by

Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.12680 v1 pith:BMDKUVF2 submitted 2025-04-17 cs.AI cs.CV

classification cs.AIcs.CV
keywords modelsreasoningembodied-rembodiedspatialtrainingcollaborativeframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans can perceive and reason about spatial relationships from sequential visual observations, such as egocentric video streams. However, how pretrained models acquire such abilities, especially high-level reasoning, remains unclear. This paper introduces Embodied-R, a collaborative framework combining large-scale Vision-Language Models (VLMs) for perception and small-scale Language Models (LMs) for reasoning. Using Reinforcement Learning (RL) with a novel reward system considering think-answer logical consistency, the model achieves slow-thinking capabilities with limited computational resources. After training on only 5k embodied video samples, Embodied-R with a 3B LM matches state-of-the-art multimodal reasoning models (OpenAI-o1, Gemini-2.5-pro) on both in-distribution and out-of-distribution embodied spatial reasoning tasks. Embodied-R also exhibits emergent thinking patterns such as systematic analysis and contextual integration. We further explore research questions including response length, training on VLM, strategies for reward design, and differences in model generalization after SFT (Supervised Fine-Tuning) and RL training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A world-model-guided MCTS pipeline synthesizes 4k web navigation trajectories and yields a WebArena Pass@3 success rate of 20.15%, above OS-Genesis (18.66%) and AgentTrek (11.94%).

  2. Reinforced Reasoning for Embodied Planning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    An SFT-plus-GRPO recipe lifts a 7B VLM to 35.6 percent success on EB-ALFRED versus 22.0 for GPT-4o-mini and 33.7 for Qwen2.5-VL-72B, with smaller but consistent gains on unseen EB-Habitat.

  3. EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent

    cs.CV 2025-07 conditional novelty 5.0 of 10

    EgoPrune prunes egomotion video tokens by homography-based frame alignment and MMR selection, keeping accuracy close to the full-token baseline while reducing compute.

  4. WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    WeThink, a 120K-image QA dataset with AI-generated reasoning paths, combined with hybrid-reward reinforcement learning, improves a 7B vision-language model across 14 benchmarks.

  5. CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CoNav lets a frozen 3D-text model pass spatial text hints to a lightly fine-tuned image-text navigation agent, improving path efficiency on several VLN benchmarks, though not all claimed state-of-the-art results hold.

  6. ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models

    cs.RO 2025-05 conditional novelty 5.0 of 10

    ManipLVM-R1 applies RLVR with IoU and trajectory-distance rewards to train a 3B VLM for affordance perception and trajectory prediction, claiming better performance and generalization than SFT on 50% of the data.

  7. Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

    cs.CL 2025-09 conditional novelty 3.0 of 10

    A survey that maps reinforcement learning methods, datasets, benchmarks, and open-source tools across the full training lifecycle of large language models, focusing on verifiable-reward reasoning.

  8. Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A thesis that combines self-learning from dialog logs, schema-guided prompting, and self-aligned factuality to build task bots with minimal human intervention.

  9. Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

    cs.CL 2025-05 conditional novelty 2.0 of 10

    A survey-style position paper claims that reinforcement fine-tuning powers reasoning in multimodal LLMs, summarizing over a hundred recent works and proposing five future research directions.

Pith tools