Pith. sign in

REVIEW 18 cited by

RVT-2: Learning Precise Manipulation from Few Demonstrations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08545 v1 pith:ARRA5IFI submitted 2024-06-12 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords rvt-2tasksdemonstrationsmanipulationeffectivefasterhighlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we study how to build a robotic system that can solve multiple 3D manipulation tasks given language instructions. To be useful in industrial and household domains, such a system should be capable of learning new tasks with few demonstrations and solving them precisely. Prior works, like PerAct and RVT, have studied this problem, however, they often struggle with tasks requiring high precision. We study how to make them more effective, precise, and fast. Using a combination of architectural and system-level improvements, we propose RVT-2, a multitask 3D manipulation model that is 6X faster in training and 2X faster in inference than its predecessor RVT. RVT-2 achieves a new state-of-the-art on RLBench, improving the success rate from 65% to 82%. RVT-2 is also effective in the real world, where it can learn tasks requiring high precision, like picking up and inserting plugs, with just 10 demonstrations. Visual results, code, and trained model are provided at: https://robotic-view-transformer-2.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

    cs.RO 2025-05 conditional novelty 7.0 of 10

    PartInstruct is a new large-scale simulated benchmark with part-level language instructions and training demonstrations; current robot policies achieve at most 31.72% average success on it.

  2. Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Aligning a VLA's latent features with instruction-selected target-object tri-views (VAE and VGGT) improves manipulation success, especially under target occlusion, with a compact 345M backbone.

  3. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  4. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  5. Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Continuous multi-view image-space keypoint trajectories plus per-camera equivariant augmentation beat strong 3D and image baselines on MimicGen and real UR5 tasks.

  6. SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation

    cs.RO 2026-03 conditional novelty 6.0 of 10

    SeedPolicy introduces self-evolving gated attention to extend the temporal horizon of diffusion policies, yielding 36.8% and 169% relative gains over standard DP on clean and randomized RoboTwin 2.0 tasks.

  7. LLaDA-VLA: Vision Language Diffusion Action Models

    cs.RO 2025-09 conditional novelty 6.0 of 10

    LLaDA-VLA applies a masked diffusion vision-language model to robot control with localized action-token classification and hierarchical decoding, achieving SOTA success rates on SimplerEnv, CALVIN, and real-robot tasks.

  8. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  9. LMPVC and Policy Bank: Adaptive voice control for industrial robots with code generating LLMs and reusable Pythonic policies

    cs.RO 2025-06 conditional novelty 6.0 of 10

    LMPVC and the Policy Bank let users control an industrial robot by voice, teach it reusable Python policies, and have a local code-generating LLM call those policies automatically.

  10. ROSA: Harnessing Robot States for Vision-Language and Action Alignment

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ROSA trains a VLA model jointly on expert actions and automatically recorded robot states, improving success rates and generalization, particularly with few demonstrations.

  11. GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    GenManip is a benchmark and simulation platform with LLM-generated scene graphs for testing how robot policies generalize to new instructions, layouts, and objects.

  12. UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    UAD distills affordance knowledge from vision-language models and DINOv2 features into a lightweight task-conditioned model that predicts pixel-level manipulation regions and improves few-shot imitation learning gener...

  13. Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    AmpAttention and RVAF raise multi-view robotic manipulation success and cut training time by suppressing attention noise with a differential-amplifier-style mechanism plus a CMRR loss.

  14. QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

    cs.CV 2025-10 conditional novelty 5.0 of 10

    Adding an auxiliary quantized-depth-token prediction task to a VLA policy improves manipulation success rates on LIBERO, Simpler, and real-robot pick-and-place tasks versus the open-pi-zero baseline.

  15. Language-Conditioned Open-Vocabulary Mobile Manipulation with Pretrained Models

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A robot system that combines GPT-4, vision-language maps, and a CLIPort-style network follows free-form household commands across rooms in simulation, reaching 10.2% average success on unseen tasks and beating two bas...

  16. RoboPearls: Editable Video Simulation for Robot Manipulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RoboPearls is a 3D Gaussian Splatting based framework that edits demonstration videos into varied photorealistic simulations, and training on them improves robot manipulation success rates on RLBench and COLOSSEUM.

  17. Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation

    cs.RO 2025-06 conditional novelty 5.0 of 10

    TUDP removes timestep conditioning from diffusion policies and adds an action-discrimination signal to learn a time-unified velocity field, achieving SOTA RLBench success rates (82.6% multi-view, 83.8% single-view) an...

  18. SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object Grasping

    cs.RO 2025-05 conditional novelty 5.0 of 10

    SR3D combines an off-the-shelf single-view 3D reconstruction model with view and keypoint matching to place the reconstructed mesh into the scene, enabling single-view grasping of transparent objects.

Pith tools