Pith. sign in

REVIEW 3 cited by

ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.02505 v2 pith:QYACGAA4 submitted 2025-03-04 cs.AI cs.CVcs.LGcs.RO

ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment

classification cs.AI cs.CVcs.LGcs.RO
keywords agenthumanrocket-2cameracross-viewgoalviewsalignment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We aim to develop a goal specification method that is semantically clear, spatially sensitive, domain-agnostic, and intuitive for human users to guide agent interactions in 3D environments. Specifically, we propose a novel cross-view goal alignment framework that allows users to specify target objects using segmentation masks from their camera views rather than the agent's observations. We highlight that behavior cloning alone fails to align the agent's behavior with human intent when the human and agent camera views differ significantly. To address this, we introduce two auxiliary objectives: cross-view consistency loss and target visibility loss, which explicitly enhance the agent's spatial reasoning ability. According to this, we develop ROCKET-2, a state-of-the-art agent trained in Minecraft, achieving an improvement in the efficiency of inference 3x to 6x compared to ROCKET-1. We show that ROCKET-2 can directly interpret goals from human camera views, enabling better human-agent interaction. Remarkably, ROCKET-2 demonstrates zero-shot generalization capabilities: despite being trained exclusively on the Minecraft dataset, it can adapt and generalize to other 3D environments like Doom, DMLab, and Unreal through a simple action space mapping.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RescueBench: Can Embodied Agents Save Lives in the Wild ?

    cs.CV 2026-06 unverdicted novelty 7.0

    RescueBench is a new diagnostic benchmark for multi-stage embodied search-and-rescue that shows no tested baselines complete the hardest tasks and identifies exploration and memory as independent failure modes.

  2. Self-supervised Hierarchical Visual Reasoning with World Model

    cs.AI 2026-05 unverdicted novelty 6.0

    ResDreamer proposes a residual-based hierarchical world model for purely self-supervised visual foresight reasoning that scales with linear communication cost and reports state-of-the-art sample and parameter efficien...

  3. Self-supervised Hierarchical Visual Reasoning with World Model

    cs.AI 2026-05 unverdicted novelty 5.0

    ResDreamer proposes a residual-reconstruction hierarchical world model for purely self-supervised visual foresight that claims SOTA sample and parameter efficiency in open-world RL.