Pith. sign in

REVIEW 4 cited by

FP3: A 3D Foundation Policy for Robotic Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08950 v1 pith:G4YM4LG2 submitted 2025-03-11 cs.RO cs.AI

classification cs.ROcs.AI
keywords foundationmodelsexistinglarge-scalemanipulationmodelobservationspolicy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Following its success in natural language processing and computer vision, foundation models that are pre-trained on large-scale multi-task datasets have also shown great potential in robotics. However, most existing robot foundation models rely solely on 2D image observations, ignoring 3D geometric information, which is essential for robots to perceive and reason about the 3D world. In this paper, we introduce FP3, a first large-scale 3D foundation policy model for robotic manipulation. FP3 builds on a scalable diffusion transformer architecture and is pre-trained on 60k trajectories with point cloud observations. With the model design and diverse pre-training data, FP3 can be efficiently fine-tuned for downstream tasks while exhibiting strong generalization capabilities. Experiments on real robots demonstrate that with only 80 demonstrations, FP3 is able to learn a new task with over 90% success rates in novel environments with unseen objects, significantly surpassing existing robot foundation models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Adding stage-wise temporal and spatial memory to a heatmap-prediction 3D VLA policy yields strong results on memory-dependent manipulation benchmarks while keeping data efficiency.

  2. See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Expressing RGB-D observations as end-effector-centered robot-frame pointmaps and adding them element-wise to RGB tokens improves pretrained VLAs under camera viewpoint variation.

  3. StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A vision-language-action model that fuses stereo-derived geometric features with semantic features improves real-world grasping success and camera-pose robustness over single-view baselines.

  4. QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

    cs.CV 2025-10 conditional novelty 5.0 of 10

    Adding an auxiliary quantized-depth-token prediction task to a VLA policy improves manipulation success rates on LIBERO, Simpler, and real-robot pick-and-place tasks versus the open-pi-zero baseline.

Pith tools