Pith. sign in

REVIEW 5 cited by

Octopus: Embodied Vision-Language Programmer from Environmental Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08588 v2 pith:GVHQH3PD submitted 2023-10-12 cs.CV cs.AIcs.LGcs.RO

classification cs.CVcs.AIcs.LGcs.RO
keywords octopusembodiedagentcodeactionexecutablefeedbackmanipulation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large vision-language models (VLMs) have achieved substantial progress in multimodal perception and reasoning. When integrated into an embodied agent, existing embodied VLM works either output detailed action sequences at the manipulation level or only provide plans at an abstract level, leaving a gap between high-level planning and real-world manipulation. To bridge this gap, we introduce Octopus, an embodied vision-language programmer that uses executable code generation as a medium to connect planning and manipulation. Octopus is designed to 1) proficiently comprehend an agent's visual and textual task objectives, 2) formulate intricate action sequences, and 3) generate executable code. To facilitate Octopus model development, we introduce OctoVerse: a suite of environments tailored for benchmarking vision-based code generators on a wide spectrum of tasks, ranging from mundane daily chores in simulators to sophisticated interactions in complex video games such as Grand Theft Auto (GTA) and Minecraft. To train Octopus, we leverage GPT-4 to control an explorative agent that generates training data, i.e., action blueprints and corresponding executable code. We also collect feedback that enables an enhanced training scheme called Reinforcement Learning with Environmental Feedback (RLEF). Through a series of experiments, we demonstrate Octopus's functionality and present compelling results, showing that the proposed RLEF refines the agent's decision-making. By open-sourcing our simulation environments, dataset, and model architecture, we aspire to ignite further innovation and foster collaborative applications within the broader embodied AI community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following

    cs.AI 2025-09 conditional novelty 6.0 of 10

    ExRAP couples LLM planning with a temporal knowledge-graph memory and information-based exploration, improving success and efficiency for continual embodied instruction following.

  2. Cracking the Code of Action: a Generative Approach to Affordances for Reinforcement Learning

    cs.AI 2025-04 conditional novelty 6.0 of 10

    A VLM-generated template-matching script masks the action space of a DQN agent, yielding large sample-efficiency gains on MiniWob++ in the low-data regime.

  3. Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Insight-V uses machine-generated long-chain reasoning data, a reasoning-plus-summary multi-agent setup, and iterative DPO to improve visual reasoning scores of multimodal LLMs.

  4. LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents

    cs.CV 2024-12 conditional novelty 5.0 of 10

    An LLM agent reads past loss weights and quality scores, then writes new loss weights, letting image processing models be trained toward non-differentiable objectives like IQA scores and text feedback.

  5. An LLM-powered Natural-to-Robotic Language Translation Framework with Correctness Guarantees

    cs.RO 2025-08 conditional novelty 4.0 of 10

    NRTrans uses a small Robot Skill Language with a compiler and iterative error feedback to improve the success rate of LLM-generated robot control programs.

Pith tools