REVIEW 4 cited by
3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present 3D FlowMatch Actor (3DFA), a 3D policy architecture for robot manipulation that combines flow matching for trajectory prediction with 3D pretrained visual scene representations for learning from demonstration. 3DFA leverages 3D relative attention between action and visual tokens during action denoising, building on prior work in 3D diffusion-based single-arm policy learning. Through a combination of flow matching and targeted system-level and architectural optimizations, 3DFA achieves over 30x faster training and inference than previous 3D diffusion-based policies, without sacrificing performance. On the bimanual PerAct2 benchmark, it establishes a new state of the art, outperforming the next-best method by an absolute margin of 41.4%. In extensive real-world evaluations, it surpasses strong baselines with up to 1000x more parameters and significantly more pretraining. In unimanual settings, it sets a new state of the art on 74 RLBench tasks by directly predicting dense end-effector trajectories, eliminating the need for motion planning. Comprehensive ablation studies underscore the importance of our design choices for both policy effectiveness and efficiency.
Forward citations
Cited by 4 Pith papers
-
Learning 3D Affordances for Blade Insertion in Cluttered Stowing
VulcanVoxel reconstructs blade occupancy with a 3D masked autoencoder, recovering multi-modal free-space affordances from unimodal warehouse stow data and raising top-5 coverage from 0.71 to 0.89.
-
ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models
A preprocessing pipeline that reconstructs a 3D point cloud from RGB images and re-renders it from a fixed viewpoint improves viewpoint robustness and data efficiency for vision-based robot policies.
-
High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching
One-step flow-matching visuomotor policy with recursive correction, dual-timestep spectral consistency, and contrastive mode separation matches or exceeds 10-step baselines at 1 NFE.
-
ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning
Co-training a diffusion visuomotor policy on unified past–current–future sparse 3D object–gripper keypoint flows improves long-horizon and non-Markovian manipulation over action-only and future-only baselines.
Discussion (0). Continue with ORCID to comment.