Pith. sign in

REVIEW 1 cited by

Multi-View Masked World Models for Visual Robotic Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.02408 v2 pith:7FXY535E submitted 2023-02-05 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords multi-viewmaskedautoencodermanipulationroboticvisualworldcameras
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual robotic manipulation research and applications often use multiple cameras, or views, to better perceive the world. How else can we utilize the richness of multi-view data? In this paper, we investigate how to learn good representations with multi-view data and utilize them for visual robotic manipulation. Specifically, we train a multi-view masked autoencoder which reconstructs pixels of randomly masked viewpoints and then learn a world model operating on the representations from the autoencoder. We demonstrate the effectiveness of our method in a range of scenarios, including multi-view control and single-view control with auxiliary cameras for representation learning. We also show that the multi-view masked autoencoder trained with multiple randomized viewpoints enables training a policy with strong viewpoint randomization and transferring the policy to solve real-robot tasks without camera calibration and an adaptation procedure. Video demonstrations are available at: https://sites.google.com/view/mv-mwm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A preprocessing pipeline that reconstructs a 3D point cloud from RGB images and re-renders it from a fixed viewpoint improves viewpoint robustness and data efficiency for vision-based robot policies.

Pith tools