REVIEW 3 cited by
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
read the original abstract
We aim to develop a goal specification method that is semantically clear, spatially sensitive, domain-agnostic, and intuitive for human users to guide agent interactions in 3D environments. Specifically, we propose a novel cross-view goal alignment framework that allows users to specify target objects using segmentation masks from their camera views rather than the agent's observations. We highlight that behavior cloning alone fails to align the agent's behavior with human intent when the human and agent camera views differ significantly. To address this, we introduce two auxiliary objectives: cross-view consistency loss and target visibility loss, which explicitly enhance the agent's spatial reasoning ability. According to this, we develop ROCKET-2, a state-of-the-art agent trained in Minecraft, achieving an improvement in the efficiency of inference 3x to 6x compared to ROCKET-1. We show that ROCKET-2 can directly interpret goals from human camera views, enabling better human-agent interaction. Remarkably, ROCKET-2 demonstrates zero-shot generalization capabilities: despite being trained exclusively on the Minecraft dataset, it can adapt and generalize to other 3D environments like Doom, DMLab, and Unreal through a simple action space mapping.
Forward citations
Cited by 3 Pith papers
-
RescueBench: Can Embodied Agents Save Lives in the Wild ?
RescueBench is a new diagnostic benchmark for multi-stage embodied search-and-rescue that shows no tested baselines complete the hardest tasks and identifies exploration and memory as independent failure modes.
-
Self-supervised Hierarchical Visual Reasoning with World Model
ResDreamer proposes a residual-based hierarchical world model for purely self-supervised visual foresight reasoning that scales with linear communication cost and reports state-of-the-art sample and parameter efficien...
-
Self-supervised Hierarchical Visual Reasoning with World Model
ResDreamer proposes a residual-reconstruction hierarchical world model for purely self-supervised visual foresight that claims SOTA sample and parameter efficiency in open-world RL.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.